Senior Data Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Lead data engineering for a petabyte-scale industrial digital twin supporting a coral restoration programme.
Build curated datasets for production AI and MLOps pipelines and own the data governance framework in a senior delivery team.
Combine hands-on delivery with advisory responsibility for data architecture and governance.
Key Skills for This Role
Full Job Posting
Role Overview
Lead data engineering for a petabyte-scale industrial digital twin supporting a coral restoration programme.
Build curated datasets for production AI and MLOps pipelines and own the data governance framework in a senior delivery team.
Combine hands-on delivery with advisory responsibility for data architecture and governance.
Databricks and Pipelines
- Build and maintain bronze, silver, and gold Databricks pipelines for curated datasets.
- Develop Databricks notebooks and jobs with PySpark and SQL, manage scheduling, and tune performance.
- Register governed datasets in Unity Catalog with metadata, ownership, and access tags.
- Apply dimensional modeling, star schema, and slowly changing dimensions to curated Lakehouse layers.
Orchestration and Ingestion
- Build Azure Data Factory pipelines and triggers for source-system ingestion into Databricks.
- Consume streaming data from Azure Event Hubs and configure trigger-driven processing.
- Develop Azure Function Apps and manage Azure Batch Accounts for downstream ingestion.
- Keep the Event Hubs, Function App, Batch Account, and Databricks chain resilient, monitored, and recoverable.
Migration, Security and Storage
- Migrate operational, imagery, scientific-archive, IoT, and sensor data while preserving integrity and metadata.
- Manage credentials and secrets in Azure Key Vault and implement RBAC and managed identities.
- Maintain Azure Data Lake storage and organize files and metadata across platform outputs.
Downstream Data Enablement
- Prepare governed datasets for Power BI, Tableau, Digital Twin applications, AI/ML, and MLOps consumption.
- Build feature stores and feature-engineering pipelines with versioned datasets, lineage, and drift or freshness monitoring.
- Build GenAI and RAG pipelines for document preparation, chunking, embeddings, vector indexes, retrieval evaluation, and access control.
- Enable graph, knowledge-graph, geospatial, and time-series data structures for analytics and decision support.
Architecture and Governance
- Own Lakehouse and application data architecture, including medallion layers, domain boundaries, data products, serving patterns, and integration standards.
- Define and implement governance covering domains, ownership, data quality, classification, privacy, metadata, retention, and lifecycle.
- Onboard domains to governance standards and direct application teams on required data quality controls.
Core Technology Skills
- Azure Databricks, Apache Spark, PySpark, Delta Lake, Unity Catalog, and medallion architecture.
- Azure Data Factory, ADLS Gen2, Event Hubs, Function Apps, and Batch Accounts.
- Advanced SQL, Python, PostgreSQL, PostGIS, and exposure to time-series stores.
- Azure Key Vault, managed identities, RBAC, Git, CI/CD, Azure DevOps, and data observability practices.
- AI/ML data engineering, feature stores, MLflow, MLOps, GenAI, RAG, embeddings, and vector stores.
Requirements
- Senior, hands-on data engineering experience delivering production data platforms and pipelines.
- Experience with Azure Databricks, Apache Spark or PySpark, Delta Lake, Unity Catalog, and medallion architecture.
- Experience building Azure Data Factory pipelines and integrating batch and streaming data sources.
- Advanced SQL and strong Python or PySpark skills for data transformation and pipeline development.
- Knowledge of dimensional modeling, star schemas, slowly changing dimensions, and data warehousing practices.
- Experience with Azure Event Hubs, Function Apps, Batch Accounts, Azure Data Lake, and trigger-based ingestion.
- Experience with data governance, ownership and stewardship, data quality, metadata, classification, retention, and lifecycle management.
- Experience implementing secure access using Azure Key Vault, managed identities, and RBAC.
- Experience preparing data for reporting, AI/ML pipelines, feature stores, and MLOps.
- Experience with Git-based development and CI/CD for data pipelines using Azure DevOps or an equivalent platform.
- Experience with geospatial, time-series, sensor, IoT, or scientific-archive data.
- Experience with GenAI and RAG data preparation, embeddings, vector indexes, retrieval evaluation, and access control.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at DigyCorp
Senior Data Engineer
, UAE
DigyCorp is seeking a Senior Data Engineer to lead delivery for a petabyte-scale industrial digital twin. The role combines hands-on Azure and Databricks engineering with data architecture, governance, streaming ingestio
Senior Product Owner
, UAE
The employer is seeking a Senior Product Owner to own the backlog and roadmap for a large Digital Twin Platform supporting a coral restoration programme. The role combines product leadership, continuous discovery, stakeh
Senior Delivery Manager
, UAE
DigyCorp is seeking a Senior Delivery Manager to lead a large technology programme supporting a national coral reef restoration initiative. The role owns roadmap, governance, budget, risk, quality, dependencies, and clie