{bc}
oracle

AWS Data Engineer

EXL
Pune, IND
Mid
Hybrid
Discovered 1 weeks ago
awsaws-glueaws-lambdadynamodbicebergkafka
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

awsaws-glueaws-lambda
Smart Apply

Full Job Posting

Responsibilities

  • Key Responsibilities
  • Design and implement scalable, secure, and cost-optimized AWS data architectures.
  • Develop and maintain ETL pipelines using AWS Lambda and AWS Glue ETL.
  • Configure and manage AWS Glue Crawlers, Glue Data Catalog, and schema evolution.
  • Build, optimize, and unit test applications on the Apache Spark framework using PySpark.
  • Design and optimize data lakes using Apache Iceberg on AWS, including table compaction and Iceberg performance tuning.
  • Work extensively with data formats such as Avro, Parquet, JSON, XML, and CSV.
  • Orchestrate event-driven workflows using AWS Step Functions and Amazon EventBridge.
  • Connect and integrate Starburst from Lambda and Glue ETL jobs for federated querying.
  • Implement CI/CD pipelines for automated testing and deployment.
  • Perform unit testing using PyTest, and performance tuning of Spark and Python applications
  • Key Responsibilities - Design and implement scalable, secure, and cost-optimized AWS data architectures. - Develop and maintain ETL pipelines using AWS Lambda and AWS Glue ETL. - Configure and manage AWS Glue Crawlers, Glue Data Catalog, and schema evolution. - Build, optimize, and unit test applications on the Apache Spark framework using PySpark. - Design and optimize data lakes using Apache Iceberg on AWS, including table compaction and Iceberg performance tuning. - Work extensively with data formats such as Avro, Parquet, JSON, XML, and CSV. - Orchestrate event-driven workflows using AWS Step Functions and Amazon EventBridge. - Connect and integrate Starburst from Lambda and Glue ETL jobs for federated querying. - Implement CI/CD pipelines for automated testing and deployment. - Perform unit testing using PyTest, and performance tuning of Spark and Python applications

Qualifications

  • Strong understanding of AWS architecture best practices, scalability, security, and cost optimization strategies.
  • Strong hands-on experience with AWS services including Lambda, Glue ETL, Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS.
  • Deep experience in Apache Spark (PySpark/Scala) development, unit testing, and performance optimization.
  • Strong Python programming skills using libraries such as pandas, requests, json, and awswrangler.
  • Experience on Apache Kafka and Confluent Kafka.
  • Experience designing and optimizing data lakes using Apache Iceberg, including compaction and Iceberg optimization techniques.
  • - Strong understanding of AWS architecture best practices, scalability, security, and cost optimization strategies. - Strong hands-on experience with AWS services including Lambda, Glue ETL, Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS. - Deep experience in Apache Spark (PySpark/Scala) development, unit testing, and performance optimization. - Strong Python programming skills using libraries such as pandas, requests, json, and awswrangler. - Experience on Apache Kafka and Confluent Kafka. - Experience designing and optimizing data lakes using Apache Iceberg, including compaction and Iceberg optimization techniques.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at EXL