{bc}
bayt

Senior Data Engineer

Unknown
Riyadh Region, KSA
Onsite
Discovered Yesterday
Data engineeringAI and ML data systemsSQLPythonPySparkAWS S3
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Data engineeringAI and ML data systemsSQL
Smart Apply

Full Job Posting

Role Overview

Own the end-to-end data layer supporting generative AI products, retrieval systems, agents, and analytics.

Build governed data systems on AWS so consumers can rely on consistent, query-ready data.

Work in a small multilingual team while owning pipelines and helping establish engineering standards.

Responsibilities

  • Build and run batch and streaming pipelines from source systems into the lake and warehouse.
  • Own raw, curated, schema, quality, and lineage layers.
  • Build retrieval components including connectors, document parsing, chunking, embeddings, and vector indexing.
  • Create curated datasets and metrics for AI and analytics consumers.
  • Add validation, monitoring, access controls, PII handling, and entitlement-aware datasets.
  • Collaborate with platform and DevOps engineers on dependable data and retrieval services.
  • Control storage, compute, query, embedding, and vector workload costs.
  • Review code, document systems, and shape data engineering standards.

Requirements

  • Eight or more years of overall data engineering experience.
  • Hands-on AI or ML data engineering experience involving retrieval, embeddings, or feature data.
  • Strong SQL and Python, including PySpark or similar distributed processing.
  • Production experience with AWS S3, Glue, Athena, and Redshift.
  • Experience with layered data architecture and raw-to-curated transformations.
  • Experience with ELT or integration tools, event-driven pipelines, streaming or CDC, semantic layers, vector stores, open table formats, and orchestration.

Technical Environment

  • AWS data stack: S3, Glue, Data Catalog, Athena, and Redshift.
  • Possible integration tools include Airbyte, Fivetran, and Meltano.
  • Possible streaming or CDC technologies include Kinesis, Amazon MSK, and Debezium.
  • Possible vector stores include pgvector, Amazon OpenSearch, Pinecone, Weaviate, and Milvus.
  • Possible orchestrators include Amazon MWAA, Step Functions, Dagster, and Prefect.
  • Open table formats include Parquet with Apache Iceberg, Delta Lake, or Hudi.

Workplace

  • The source metadata classifies the role as onsite in the Riyadh Region.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Unknown