Pipeline Engineering: Build and maintain high-throughput ETL/ELT pipelines that ingest data from various sources into our Lakehouse.
Code Quality & Tooling: Drive the adoption of dbt for transformation and PySpark for heavy lifting. You will be responsible for writing modular, reusable code that follows strict CI/CD practices.
Observability & Reliability: Implement "Data SLAs." You will build the monitoring and alerting systems that notify the team of data drift or pipeline failures before the business notices.
Data Productization: Work closely with the BI team and Data Scientists to prepare "feature-ready" datasets.
Performance Tuning: Deep-dive into SQL and Spark query plans to optimize slow-running jobs, reducing cost and latency.
Local Development Advocacy: Implementing a workflow where engineers can develop and test complex SQL transformations locally using DuckDB, drastically reducing "waiting-for-cluster" time and lowering development costs.
Cost-Efficient Micro-Pipelines: Building specialized pipelines for specific client reporting or data-quality checks that run on single-node containers (e.g., Lambda or small ECS tasks) using DuckDB, avoiding the minimum billing cycles of larger warehouses.
Embedded Analytics: Exploring ways to use DuckDB as an embedded engine for internal tools or "edge" processing within the BPO's local site offices where bandwidth to the cloud might be a constraint.
5+ Years in Data Engineering: You've lived through the "on-call" life and know how to build systems that don't break at 3 AM.
The Power Trio: Expert-level proficiency in Python , SQL , and PySpark .
Modern Lakehouse Stack: Hands-on experience with Databricks (Delta Lake) or Snowflake . You understand the nuances of the Medallion Architecture (Bronze/Silver/Gold).
Transformation & Modeling: Advanced experience with dbt (Data Build Tool) . You treat data models like software, including version control, testing, and documentation.
Orchestration: Experience with Apache Airflow or Prefect, specifically building complex, idempotent DAGs.
Streaming Experience: Familiarity with Kafka , Kinesis , or Spark Streaming is a huge plus-our BPO operations move in real-time.
Infrastructure as Code (IaC): Comfort with Terraform or CloudFormation to manage your own data infrastructure.
In-Process Analytics (DuckDB): Proven experience using DuckDB for high-speed local development, unit testing data transformations, or as a query engine for "small-to-medium" datasets (up to 100GB) without the overhead of a distributed cluster.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
TaskUs is seeking a Team Leader to manage day-to-day team operations and deliver strong customer experiences. The role focuses on coaching teammates, meeting SLAs and KPIs, maintaining schedules, coordinating with other
Hybrid Execution Patterns: Ability to identify when to use heavyweight compute (PySpark/Databricks) versus lightweight compute (DuckDB/Python) to minimize cloud costs and reduce job latency.
Parquet/Iceberg Interaction: Experience using DuckDB to directly query data stored in S3/Azure Blob (via Parquet or Iceberg files) for rapid ad-hoc analysis or local dashboarding.
Orchestration: Kubernetes (K8s) . Deep understanding of Pods, Deployments, Services, ConfigMaps, and Secrets management.
TaskUs is proud to be an equal opportunity workplace and is an affirmative action employer.
We celebrate and support diversity; we are committed to creating an inclusive environment for all employees.
TaskUs people first culture thrives on it for the benefit of our employees, our clients, our services, and our community.
About TaskUs
Business Services47000 employeesFounded 2008
Provides outsourced digital business and customer experience services.