zoho_recruit
Sr. Data Engineer - Spark
Techsa
Remote, USA
Full-time
Senior · 7+ years experience
Remote
Discovered 2 days ago
Apache SparkJavaScalaPythonSpark Structured StreamingYARN
Free
Job Fit Check
Base Career helps you apply smarter for this job.
?%
Ready to ScanKey skills for this role
Apache SparkJavaScala
Key Skills for This Role
Apache SparkJavaScalaPythonSpark Structured StreamingYARN
Full Job Posting
Requirements
- 7+ years of experience in data engineering and software development
- Ability to write high-quality code in Java/Scala, Python, or equivalent languages
- Deep, hands-on production experience with Apache Spark — batch and Spark Structured Streaming (core requirement)
- Demonstrated Spark performance tuning: partitioning, caching and persistence, broadcast joins, shuffle reduction, data-skew handling, and Adaptive Query Execution
- Experience operating Spark on self-managed clusters (YARN, Kubernetes, or standalone) — executor sizing, resource allocation, and multi-tenant workloads
- Practical experience with Kafka (or equivalent messaging systems) as a Spark source and sink for high-volume workloads, including offset and checkpoint management
- Practical experience with distributed query engines (e.g., Trino/Presto or similar)
- Practical experience with ETL / data integration tools, commercial or open-source (e.g., Datastage, Informatica, Apache NiFi, or similar)
- Practical experience with SQL-based transformation frameworks (e.g., dbt or others)
- Strong SQL skills and understanding of data modeling and data warehousing for analytical workloads
- Hands-on experience with real-time / low-latency analytical stores (columnar or OLAP engines, e.g., Apache Pinot/ClickHouse or similar)
- Practical experience with big-data platforms and distributions (e.g., Cloudera, Hadoop ecosystem, Databricks, or similar)
- Practical experience containerizing and operating data workloads (Docker; Kubernetes a plus)
- Experience with workflow orchestration tools (e.g., Airflow or similar)
- Familiarity with data lake table formats (e.g., Apache Iceberg, Delta Lake, or similar), including schema evolution and compaction
- Familiarity with data governance / cataloging tools (e.g., DataHub or similar)
- Familiarity with lakehouse management systems (e.g., Apache Amoro or similar)
- Familiarity using AI tools for development and debugging (Claude, Cursor, Codex)
About Techsa
Software Development25 employeesFounded 2022
Provider of software solutions and digital transformation services.
Visit company websiteJobs and hiring trendsApply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer