{bc}
oracle

Data Engineer

EXL
Bengaluru, IND
Mid · 5–8 years experience
Hybrid
Discovered 1 weeks ago
awsdatabricksdelta-lakehadooplangchainlooker
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

awsdatabricksdelta-lake
Smart Apply

Full Job Posting

Responsibilities

  • ROLES AND RESPONSIBILITIES: • Cloud Modernization & Migration: Deconstruct legacy Apache Spark and Hadoop MapReduce workflows to re-architect and rebuild them as optimized, production-ready pipelines within AWS and Databricks. • AI Agent & LLM Development: Design, build, and deploy intelligent AI agents and workflow automation tools leveraging leading Large Language Models (LLMs) such as Anthropic Claude. • AI Data Pipeline Engineering: Build and optimize pipeline architectures explicitly tailored for AI use cases, including unstructured data ingestion, real-time feature tokenization, and metadata tagging for vector databases. • Legacy Infrastructure Maintenance: Monitor, maintain, and troubleshoot existing big data workloads running on our Hadoop cluster to guarantee data availability for business operations during the multi-phase migration, resolving bottlenecks and Out Of-Memory (OOM) errors. • BI Engineering & Support: Act as the primary engineering liaison for downstream business stakeholders utilizing BI tools (e.g. Tableau, Looker, etc.) by performing minor functional enhancements, bug fixes, and data extract optimizations to resolve report dashboard latency. • Cloud Optimization: Utilize Databricks and Delta Lake features (e.g., ACID transactions, Z-Ordering, caching) to significantly improve pipeline performance, reliability, and cost efficiency. • Databricks AI Suite Implementation: Leverage Databricks tools (such as Databricks Vector Search, Mosaic AI, and Lakeflow) to orchestrate, track, and serve productio

Qualifications

  • AI & Agentic Frameworks: Hands-on experience or deep technical familiarity building functional AI agents, integrating LLM APIs (specifically Anthropic Claude), and utilizing orchestration frameworks (e.g., LangChain, Databricks Mosaic AI Agent Framework or any other tool ). • Distributed Computing: Foundational understanding of distributed storage and computing concepts—specifically partitioning, shuffling, caching, and broadcast joins. Solid hands-on experience with Apache Spark is required. • Programming & SQL: Strong proficiency in Python (PySpark) or Scala, alongside intermediate-to-advanced SQL querying capabilities (window functions, query tuning, and complex joins). • Cloud & Databricks Exposure: Direct experience or deep theoretical knowledge of the AWS ecosystem (S3, IAM) and Databricks environments. • Visualization Layer: Practical experience working with any BI tool (e.g. Tableau, Looker, Power BI, etc.) with the capability to debug calculated fields, modify parameters, and troubleshoot slow-loading reports.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at EXL