•Cloud Infrastructure Provisioning & Management
•Provision, configure, and manage core cloud infrastructure across AWS and Azure supporting Irth’s Databricks-based data estate.
•Manage infrastructure components including:
•VPCs / VNets
•Subnets
•Route tables and network connectivity
•S3 buckets / Azure Storage Accounts
•Compute resources
•Private endpoints and related networking services
•Establish and maintain secure DEV, QA, and PROD environments with appropriate isolation and platform guardrails.
•Support cloud landing-zone implementation and ensure environments follow established security, governance, and cost policies.
•Use Infrastructure-as-Code (IaC) technologies such as:
•Terraform
•Azure ARM/Bicep
•AWS CloudFormation
•Version, review, test, and deploy infrastructure changes through controlled engineering workflows.
•Troubleshoot infrastructure, networking, and cloud-service issues affecting data and ML workloads.
•Identity & Access Management (IAM)
•Provision and manage IAM roles, service principals, managed identities, groups, and access policies across AWS, Azure, and Databricks.
•Implement least-privilege access and role-based access control (RBAC) aligned with enterprise security standards.
•Configure and support identity federation and automated identity lifecycle management using technologies such as:
•SSO
•SCIM
•SAML
•OAuth
•Manage access across cloud resources, Databricks workspaces, Unity Catalog, storage, compute, and other platform services.
•Review and process elevated or administrative access requests through controlled and auditable procedures.
•Balance engineering-team productivity with security, compliance, and least-privilege requirements.
•Participate in periodic access reviews and remediation of excessive or inappropriate permissions.
•Databricks Platform Administration & Team Onboarding
•Create, configure, and administer Databricks workspaces and associated platform resources.
•Configure and manage Unity Catalog metastores, catalogs, schemas, storage credentials, external locations, and permissions .
•Support onboarding of Data Engineering, Data Science, and MLOps projects by establishing the appropriate platform resources and access.
•Troubleshoot onboarding blockers, including:
•Missing administrative privileges
•Cluster-policy conflicts
•Workspace configuration issues
•Catalog and schema permission problems
•Storage-access failures
•Identity and authentication issues
•Define and maintain Databricks cluster policies to control compute configurations, security, and cost.
•Configure instance profiles, managed identities, storage credentials, and other mechanisms required for secure data access.
•Establish workspace-level guardrails for compute, networking, storage, and user access.
•Partner with Data Engineering, Data Science, and MLOps teams to ensure new projects can be onboarded efficiently with the appropriate catalogs, schemas, permissions, and infrastructure already available.
•Act as the primary escalation point for Databricks platform and infrastructure issues.
•Security, Networking & Compliance
•Implement secure private connectivity between Databricks, cloud services, and enterprise resources.
•Configure and troubleshoot technologies such as:
•AWS PrivateLink / VPC endpoints
•Azure Private Link / private endpoints
•VNet/VPC peering
•Secure cluster connectivity
•Network security groups and firewall controls
•Egress restrictions
•Implement encryption and key-management practices using:
•AWS KMS
•AWS Secrets Manager
•Azure Key Vault
•Equivalent enterprise security services
•Establish appropriate network segmentation and isolation between environments.
•Support security reviews and enterprise compliance requirements.
•Maintain audit logging and provide evidence for:
•Access reviews
•Administrative activity
•Infrastructure changes
•Data access
•Retention controls
•Security and governance reviews
•Help ensure infrastructure and platform configurations align with applicable regulatory and organizational requirements.
•Cost Management, Automation & Operations
•Implement and enforce cloud and Databricks tagging standards across infrastructure and workloads.
•Monitor infrastructure and Databricks spend and identify opportunities for cost optimization.
•Support showback/chargeback initiatives and cost reporting.
•Implement resource policies and controls that prevent unnecessary or unapproved cloud and Databricks spend.
•Automate recurring infrastructure and access-management activities, including:
•Workspace provisioning
•Catalog and schema setup
•IAM role assignment
•Access provisioning
•Environment configuration
•Standard platform onboarding
•Build reusable automation that reduces manual provisioning and improves onboarding turnaround time.
•Monitor platform health and respond to infrastructure and access-related incidents.
•Troubleshoot cloud, networking, identity, and Databricks issues affecting production workloads.
•Maintain clear operational runbooks, architecture documentation, access procedures, troubleshooting guides, and onboarding documentation .
Role Outcomes
Success in this role means that Irth’s data and ML teams can rely on a secure, available, well-governed, and easy-to-use platform .
Key outcomes include:
•Fast and repeatable provisioning of cloud and Databricks infrastructure.
•Efficient onboarding of Data Engineering, Data Science, and MLOps teams.
•Secure, auditable, least-privilege access across AWS, Azure, and Databricks.
•Stable and well-governed DEV, QA, and PROD environments.
•Reduced infrastructure and platform-related blockers for engineering teams.
•Increased automation of provisioning and access-management processes.
•Strong adherence to networking, security, governance, and compliance standards.
•Improved visibility and control over cloud and Databricks costs.
•Well-documented infrastructure and operational procedures.
•A platform that enables engineering teams to build and ship without unnecessary infrastructure friction .