Base Career helps you apply smarter for this job.
Key skills for this role
Build and maintain modular Terraform components that provision cloud and Databricks platform capabilities, including workspaces, cluster policies, secret scopes, networking, and Unity Catalog storage credentials and external locations, consistently across environments and clouds.
Design and operate automated GitLab CI/CD pipelines that deploy data platform code and infrastructure through development, staging, and production with approval gates and controlled promotion.
Monitor data quality and pipeline health, own first-line response to production alerts, resolve issues directly where possible, and escalate only when the root cause requires deeper pipeline or platform change.
Build, test, and continuously improve runbooks for recurring failure modes, automating remediation when the response can be made safe and repeatable.
Apply AI- and agent-assisted operations capabilities to accelerate detection, triage, and remediation as those capabilities mature.
Build end-to-end observability across Databricks, Snowflake, and orchestration services by bringing platform and execution metadata into unified dashboards, alerting, and pipeline service-level monitoring.
Containerize and operate platform workloads, orchestration, and CI runners on Docker and Kubernetes across EKS and AKS.
Operate and re-host Apache Airflow, consolidating self-managed instances onto Kubernetes or a managed runtime, while supporting migration of legacy build tooling to GitLab CI.
Embed security operations into the platform lifecycle through secrets management, short-lived credentials, least-privilege access, encryption key policies, audit logging, and vulnerability remediation.
Drive cloud cost optimization through cost attribution, policy enforcement, right-sizing, tagging standards, and anomaly alerting.
Support secure cross-cloud connectivity and data movement between AWS-resident source systems and the Azure data platform.
Identify opportunities for automation, reusable standards, reliability improvement, and technical debt reduction, and provide technical guidance to team members.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Burgess Hill, GBR
Bristol, GBR
Bristol, GBR
Burgess Hill, GBR
Ahmedabad, IND
Pune, IND
Eastbourne, GBR
, IND
Clevedon, GBR
Bachelor's Degree in Computer Science , Engineering, Information Technology, or a related field and a minimum of 5 years of relevant experience in DevOps, platform engineering, site reliability engineering, or cloud infrastructure roles.
Strong hands-on experience administering and automating workloads on both Amazon Web Services and Microsoft Azure, including identity, networking, and storage services.
Demonstrated expertise building modular, reusable infrastructure with Terraform across multiple environments.
Proven experience designing and maintaining CI/CD pipelines in GitLab CI or an equivalent platform, including environment promotion, approval gates, and secrets handling.
Strong proficiency with Docker and Kubernetes, including EKS or AKS and running CI/CD or orchestration workloads on Kubernetes.
Hands-on experience operating Apache Airflow, including DAG deployment, scheduling, retries, backfills, pools, and failure recovery.
Solid Linux systems administration, shell scripting, and troubleshooting skills.
Hands-on experience implementing metrics, logging, dashboards, and alerting using tools such as Datadog or Grafana.
Demonstrated experience monitoring pipeline and data quality, owning first-line triage of production alerts, and writing and executing runbooks for recurring failures.
Working knowledge of secrets management, least-privilege access control, encryption key management, audit logging, and vulnerability remediation.
Working knowledge of cloud cost monitoring, attribution, tagging standards, and right-sizing practices.
Proficiency in Python or another scripting language used for automation and tooling.
Excellent analytical, troubleshooting, communication, and organizational skills with the ability to work across cross-functional and geographically distributed teams.
Ability to manage competing priorities in a fast-paced environment while maintaining attention to quality, controls, and long-term outcomes.
Adhere to all company rules and requirements, including Environmental Health & Safety rules, and take adequate control measures to prevent injuries, protect the environment, and prevent pollution under their span of influence or control.
Experience operating Databricks or Snowflake, including platform administration, policy management, and service principal or access management.
Experience with Databricks Asset Bundles, Unity Catalog, the Databricks Terraform provider, or managed Terraform state and gated production applies.
Experience with keyless authentication from CI, such as OIDC-based role assumption or IRSA.
Experience with data observability or AI-assisted operations capabilities, such as Monte Carlo, Datadog AI capabilities, or Databricks operations automation.
Experience with Astronomer, MWAA, or Airflow on Kubernetes using KubernetesExecutor .
Experience with Azure Data Factory, ADLS Gen2, Key Vault, Azure Monitor, and Entra ID, or AWS Glue, Athena, Lambda, and Secrets Manager.
Experience supporting hybrid or cross-cloud architectures and private connectivity among clouds and on-premises environments.
Experience migrating CI/CD from Jenkins or another legacy build platform.
Experience with streaming or change data capture technologies such as AWS DMS, Amazon MSK, or Kafka.
Experience with SALT, Ansible, Azure DevOps, or related configuration and delivery tooling.
Relevant experience in medical device , pharmaceuticals, or another highly regulated environment, including validation and change control expectations.
Relevant AWS, Azure, Kubernetes, or Terraform certifications.
Develops solutions to a variety of complex platform and operational problems and exercises judgment in selecting methods and techniques after considering risk and alternatives.
Works with general direction focused on end results and may provide technical guidance to lower-level personnel.
Participation in an on-call rotation may be required to support platform reliability.
Travel may be required up to 5%.
Verified company details for this employer are not available yet.
Full-time
Senior · 5+ years experience
Onsite
Apply faster on company sites with our extension.