On-Premises Data Protection Platform Administration
•Administer the data protection estate, including Rubrik, Veeam and NetBackup, covering policy configuration, protection groups, SLA domains, job scheduling and agent or connector health.
•Administer tape infrastructure, including libraries, drives, media pools, barcode and slot management, and offsite media rotation and vaulting.
•Monitor backup and replication job outcomes daily, investigate failures through to root cause, and remediate rather than re-run without diagnosis.
•Maintain immutable and air-gapped protection copies for ransomware resilience and verify that immutability is enforced rather than merely configured.
•Maintain platform currency, including software and firmware upgrades, and manage vendor cases through to closure.
•Ensure every protected workload is covered by an appropriate policy, and identify & onboard workloads that are unprotected, across both the on-premises and cloud estates.
Cloud Data Protection – AWS and Rubrik Cloud
•Administer backup and archive to Amazon Web Services, including S3 and Glacier storage classes, bucket and lifecycle policy, versioning, retention and expiry, and transition rules between tiers.
•Configure and maintain immutability in cloud, including S3 Object Lock and AWS Backup Vault Lock, and verify that WORM retention is enforced rather than merely configured.
•Administer AWS-native protection where used, including AWS Backup, EBS and EC2 snapshots, RDS snapshots and EFS or FSx backups, and ensure cloud workloads are covered by policy on the same basis as on-premises workloads.
•Manage replication from the primary on-premises Rubrik clusters to Rubrik Cloud, including archival location configuration, replication and retention policy, seeding, bandwidth management and reconciliation of what has actually landed against what should have.
•Maintain the IAM roles, policies and KMS keys that backup and restore depend on, applying least privilege and cross-account separation so that a compromise of the production account cannot destroy the protection copies.
•Maintain connectivity and throughput for backup traffic, including VPC endpoints, Direct Connect or VPN capacity, and manage the impact of backup and restore windows on shared bandwidth.
•Execute and test recovery from cloud, including restore from Glacier tiers with retrieval time and cost understood in advance, recovery of on-premises workloads into AWS, and recovery of cloud-native workloads.
•Monitor cloud protection cost, including storage class placement, retrieval and egress charges, orphaned snapshots and unattached volumes, and report variance against budget.
Backup, Restore and Recovery Operations
•Execute restore and recovery requests across files, virtual machine, database and application workloads, meeting agreed recovery targets.
•Act as second and third level support for complex data protection, storage and compute recovery issues, troubleshooting through to root cause across backup agent, array, hypervisor, database and network layers.
•Perform data copy and migration activity between arrays, sites, tiers and cloud targets, including cutover planning, integrity validation and rollback.
•Support legal hold, eDiscovery and audit requests requiring retrieval of retained or archived data.
•Participate in priority and major incidents where data loss, corruption or platform failure is involved, and provide recovery estimates that can be relied on.
Restoration Testing Cadence and Service Coordination
•Operate a defined cadence of restoration tests across every workload type, covering file, virtual machine, database and application-level recovery, so that each type is proven within the agreed interval rather than on request.
•Agree the test schedule, scope and acceptance criteria with the infrastructure teams and the respective service and application owners and secure their participation in validation of the restored data.
•Document the outcome of every test, including failures, the corrective action taken and the revised recovery estimate, and re-test after remediation.
•Publish restoration test results to IT leadership, information security, audit and the service owners, and hold outstanding findings visible until closed.
•Maintain an evidence pack sufficient for audit, so that recoverability can be demonstrated without ad hoc reconstruction.
Disaster Recovery, RTO/RPO Adherence and DR Drills
•Maintain and improve disaster recovery capability for in-scope systems, including replication topology, failover design and runbook maintenance.
•Maintain the per-workload recovery time and recovery point position against the Institute's global RTO and RPO standards and escalate formally where the achievable position falls short of the stated requirement.
•Participate actively in disaster recovery drills as an executing engineer, covering failover, failback and validation, rather than observing or reporting on them.
•Plan and execute DR tests with the infrastructure teams and service owners, report results with remediation actions and revised recovery estimates, and track findings to verified closure.
•Author, maintain and improve runbooks for backup, restore, failover and recovery so that work is repeatable, auditable and transferable across shifts and across the team.
•Follow and enforce change control, including rollback planning and coordination of maintenance, drill and test windows with application owners.
Storage, Compute and Capacity Management
•Administer enterprise SAN and NAS storage, including volume and LUN provisioning, snapshots, array-based replication, zoning support and performance troubleshooting.
•Administer the supporting compute and virtualization platforms as they relate to protection and recovery, including VMware and Nutanix hosts and clusters, Windows and Linux workloads, and the database backup interfaces for SQL Server and Oracle.
•Maintain hands-on command of the physical estate, including racking, cabling, firmware levels and OEM support status for storage, tape and server hardware.
•Manage capacity planning and forecasting across the primary storage, compute and protection estates, and raise procurement requirements ahead of constraint rather than at it.
•Optimize cost and efficiency across both estates through retention rationalization, deduplication and compression, tiering and archive placement, cloud storage class placement, and removal of orphaned snapshots and unattached volumes.
•Manage the lifecycle of storage, compute and backup hardware, media and licensing, including refresh planning, media retirement and secure disposal.
Reporting, Automation, Documentation and Shift Handover
•Produce and maintain regular backup, restore, restoration-test, capacity and compliance reporting on an agreed cadence for IT leadership, information security, audit and service owners.
•Develop scripting and automation using PowerShell, Python or Bash, and platform APIs, to eliminate repetitive administration and to automate protection reporting.
•Maintain runbooks, procedures, configuration and retention documentation, and provide technical guidance and mentoring to less experienced engineers.
•Work the assigned shift within the infrastructure operating model, reporting to the Manager IT Infrastructure for that shift.
•Complete a documented handover at the end of each shift covering running jobs, failed backups, open restores, in-flight migrations, active incidents and pending actions, and verify the handover received at shift start by confirming ownership of every open item.
•Coordinate with the other shifts and with infrastructure teams so that protection and recovery activity continues without loss of context across regions.
JOB COMPETENCIES (Skills & Abilities)
•Technical depth: Advanced, practical administration of enterprise data protection platforms, including Rubrik, Veeam and NetBackup, and of tape infrastructure and media management.
•Storage and compute capability: Hands-on command of enterprise SAN and NAS storage, and of the supporting compute and virtualization platforms, sufficient to provision, replicate, protect and troubleshoot across the stack.
•Cloud data protection: Practical command of AWS as a backup and archive target — S3 and Glacier storage classes, lifecycle and retention policy, Object Lock and Vault Lock immutability, AWS Backup and native snapshots — together with Rubrik Cloud replication, archival configuration and cloud recovery.
•Hybrid judgement: Understands where a workload should be protected and to which tier, and can weigh recovery speed, retention obligation, immutability and cost across on-premises, AWS and Rubrik Cloud rather than defaulting to one.
•Cloud security and access discipline: Applies least privilege through IAM roles and policies, manages KMS keys and cross-account separation, and treats the protection copies as the asset an attacker will target first.
•Recovery orientation: Treats recoverability as something proven by test rather than inferred from a backup job that reports success and pursues restore failures with the same urgency as outages.
•RTO and RPO command: Knows the recovery position of every in-scope workload against the global standard, measures it rather than estimates it, and escalates a shortfall plainly rather than carrying it quietly.
•Drill discipline: Prepares for and executes disaster recovery drills as an operator, validates the result honestly, and treats a failed drill as the point of the exercise rather than an embarrassment.
•Reporting cadence: Produces protection, restoration-test and capacity reporting on schedule and without manual assembly, so that leadership, security and service owners see the same picture.
•Stakeholder coordination: Works with infrastructure teams and service and application owners to agree test scope, secure validation and close findings, rather than testing in isolation.
•Capacity and cost discipline: Forecasts growth across primary, computing and protection estates including cloud, and optimizes retention, deduplication, storage class placement and tiering so that protection cost, including retrieval and egress, stays proportionate to the risk it removes.
•Operational discipline: Runbook-driven execution, adherence to change control, and structured validation before and after every change, migration or test.
•Automation mindset: Uses scripting and platform APIs to remove repetitive administration and to automate protection reporting, treating repeated manual work as a defect.
•Problem solving: Strong analytical skills with proven ability to identify root cause across backup agent, array, hypervisor, database and network layers.
•Risk judgement: Distinguishes a protection gap that must be closed immediately from one that can be documented and scheduled and can defend the distinction.
•Shift and handover discipline: Works effectively within a shift-based operating model and transfers ownership of running jobs, open restores and in-flight migrations explicitly at shift boundaries.
•Documentation and knowledge transfer: Clear, maintainable runbooks and retention documentation that reduce single-person dependency and satisfy audit.
•Communication: Excellent written and verbal communication in English, including the ability to explain a recovery estimate, a failed test or a protection gap plainly to application owners and business stakeholders.
•Accountability and autonomy: Manages the protection estate and makes sound technical decisions with limited information, including under recovery pressure.
•Leadership and mentoring: Provides guidance and technical leadership to less experienced engineers without holding formal authority.
MINIMUM QUALIFICATIONS (Knowledge & Experience)
•Bachelor’s degree in computer science, Information Technology, Engineering or a closely related field, or equivalent professional experience. (Required)
•7+ years of experience in systems engineering, storage administration or data protection, including substantial hands-on responsibility for a production backup estate. (Required)
•Demonstrated hands-on administration of at least two enterprise data protection platforms from among Rubrik, Veeam, NetBackup or comparable products, including policy design and failure remediation. (Required)
•Demonstrated hands-on administration of enterprise SAN and NAS storage, including volume and LUN provisioning, snapshots, replication and capacity management. (Required)
•Demonstrated hands-on experience with compute and virtualization platforms as they relate to protection and recovery, including VMware or Nutanix, Windows and Linux workloads, and database backup interfaces such as SQL Server or Oracle. (Required)
•Demonstrated experience with tape infrastructure, including libraries, media management and offsite rotation. (Required)
•Demonstrated experience executing restores and recoveries across files, virtual machine, database and application workloads. (Required)
•Demonstrated experience operating a scheduled restoration testing program, including test design, execution with service owners, evidence capture and remediation of findings. (Required)
•Demonstrated experience maintaining per-workload recovery time and recovery point objectives against a defined organizational standard, including measurement and escalation of shortfalls. (Required)
•Demonstrated active participation in disaster recovery drills as an executing engineer, including failover, failback and validation. (Required)
•Demonstrated experience producing recurring backup, recovery, capacity and compliance reporting for IT leadership, security or audit. (Required)
•Demonstrated hands-on experience using AWS as a backup and archive target, including S3 and Glacier storage classes, bucket and lifecycle policy, versioning, retention and cross-region or cross-account copies. (Required)
•Demonstrated hands-on experience with immutability in cloud, including S3 Object Lock or AWS Backup Vault Lock, and with IAM roles, policies and KMS encryption as they apply to backup access and recovery. (Required)
•Demonstrated hands-on experience with Rubrik replication and archival to Rubrik Cloud, or with equivalent vendor cloud replication and archive from an on-premises backup platform. (Required)
•Demonstrated experience protecting AWS-native workloads, including AWS Backup, EBS and EC2 snapshots, and RDS, EFS or FSx backups. (Required)
•Demonstrated experience executing recovery from cloud, including restore from archive tiers with retrieval time and cost understood in advance. (Required)
•Experience managing cloud protection cost, including storage class placement, retrieval and egress charges, and removal of orphaned snapshots. (Required)
•Experience working within a shift-based infrastructure team, including structured handover of running jobs and open recovery work. (Required)
•Scripting proficiency in PowerShell, Python or Bash, with practical use of platform APIs for reporting and bulk administration. (Required)
•Professional proficiency in spoken and written English, sufficient to coordinate with global infrastructure and service teams and to communicate with senior stakeholders. (Required)
•Certification in a relevant platform such as Rubrik, Veeam Certified Engineer (VMCE), NetBackup, or a storage vendor credential from NetApp, Dell, Pure Storage or equivalent. (Preferred)
•AWS certification such as AWS Certified Solutions Architect – Associate or AWS Certified SysOps Administrator. (Preferred)
•ITIL 4 Foundation, or equivalent demonstrated knowledge of incident, change and problem practice. (Preferred)
Disclaimer: This job description indicates in general terms, the type and level of work performed as well as the typical responsibilities of employees in this classification and it may be changed by management at any time. Other duties may also apply. Nothing in this job description changes the at-will employment relationship existing between the Company and its employees.