Architect, build, and evolve a scalable enterprise data warehouse on Amazon Redshift, applying industry-standard concepts including star schemas, snowflake schemas, normalization, denormalization, referential integrity, and performance optimization strategies
Design and implement bronze, silver, and gold data layer architecture (medallion architecture): raw ingestion, cleansed and standardized intermediate layers, and curated, business-ready data products optimized for analytics consumption
Develop dimensional data models, fact and dimension tables, slowly changing dimensions (SCDs), and aggregate structures that support BI tooling, ad-hoc analytics, and downstream API consumption
Apply rigorous data modeling practices including schema design, constraint definition, indexing strategy, sort keys, distribution keys, and query plan optimization within Redshift and connected systems
Build, own, and maintain robust batch and streaming data pipelines that ingest data from disparate source systems including REST APIs, flat files, IBM DB2, MySQL, Amazon Aurora, Amazon DynamoDB, and PostgreSQL
Implement real-time and near real-time data streaming architectures using AWS-native services such as Kinesis Data Streams, Kinesis Firehose, MSK (Managed Kafka), and EventBridge to support low-latency data delivery requirements
Design pipeline frameworks for data extraction, transformation, and loading (ETL/ELT) using tools such as AWS Glue, dbt, Apache Airflow, or equivalent orchestration platforms
Ensure pipeline reliability, idempotency, fault tolerance, and automated recovery; build alerting and observability into every data workflow from day one
Own data quality end-to-end: design and implement automated profiling, cleansing, deduplication, standardization, and validation frameworks that enforce data integrity at each layer of the medallion architecture
Build and continuously evolve tooling and processes to support data governance including data cataloging, lineage tracking, metadata management, access controls, and data classification
Define and enforce data contracts between source systems and the warehouse, establishing clear SLAs for freshness, completeness, and accuracy
Partner with data consumers — Data scientists, Data analytics engineers, BI developers, product managers, and external API clients — to understand consumption patterns and ensure data products meet quality and performance expectations
Support the architecture and buildout of a reliable, scalable Data as a Service (DaaS) product, enabling external and internal consumers to access curated Transflo data via governed APIs and data sharing mechanisms
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
Contribute to the data platform infrastructure using infrastructure-as-code practices (Terraform), ensuring all data infrastructure is version-controlled, reproducible, and auditable
Design for scale: apply partitioning strategies, workload management (WLM) tuning, concurrency scaling, and caching patterns to sustain performance under high-traffic analytical and operational workloads
Champion security and compliance best practices across the data platform: column-level security, row-level access controls, encryption, and audit logging
Collaborate with software engineers, mobile platform teams, and DevOps to ensure upstream application data is well-structured, well-documented, and reliably delivered to the data platform
Leverage AI-assisted development practices and tooling to accelerate pipeline development, automate data quality checks, and improve engineering velocity
5+ years of professional data engineering experience with a track record of building and operating production-grade data warehouses and pipeline infrastructure
Expert-level experience with Amazon Redshift including cluster sizing, WLM configuration, distribution and sort key optimization, vacuuming, and query plan analysis
Deep proficiency in SQL for complex analytical queries, window functions, CTEs, stored procedures, and performance tuning across Redshift and ANSI-compatible engines
Hands-on experience ingesting data from heterogeneous source systems: REST APIs, IBM DB2, MySQL, Amazon Aurora (MySQL and PostgreSQL-compatible), Amazon DynamoDB, PostgreSQL, and file-based sources (CSV, JSON, Parquet, Avro)
Proven experience designing and implementing medallion (bronze/silver/gold) or equivalent layered data architectures at enterprise scale
Strong working knowledge of star schema and snowflake schema design, dimensional modeling theory, slowly changing dimensions, and fact table granularity decisions
Experience building real-time or near real-time data pipelines using streaming technologies such as Amazon Kinesis, Apache Kafka (or Amazon MSK), or equivalent
Proficiency with ETL/ELT orchestration tools such as AWS Glue, dbt, Apache Airflow, or AWS Step Functions
Demonstrated experience implementing data governance practices: data catalogs (AWS Glue Data Catalog, Apache Atlas, or equivalent), lineage, metadata tagging, and access control frameworks
Infrastructure-as-code experience with Terraform for provisioning and managing data infrastructure on AWS
Strong Python skills for pipeline development, data transformation logic, and automation scripting
Deep understanding of data reliability engineering: idempotency, exactly-once processing, late-arriving data handling, schema evolution, and SLA-driven pipeline design
Experience in the transportation, logistics, trucking, or fleet management industry, or with high-volume transactional SaaS platforms processing operational telemetry data is a huge plus
Experience building DaaS or data product offerings including governed external data APIs, Redshift Data Sharing, or AWS Data Exchange integrations
Knowledge of columnar storage formats (Parquet, ORC) and lakehouse patterns using Amazon S3 as a data lake layer in conjunction with Redshift Spectrum or AWS Glue
Familiarity with BI and analytics consumption tools such as Tableau, Power BI, Amazon QuickSight, or Looker and how data model design decisions impact end-user query performance
Experience with data observability platforms such as Monte Carlo, Great Expectations, or dbt tests for automated data quality monitoring
Contributions to reusable data platform tooling, shared dbt packages, or internal data engineering frameworks
Experience working in fully remote, distributed engineering teams
About Transflo
Software & SaaS302 employeesFounded 1991
Private transportation-software company serving carriers, brokers, factors, shippers, and drivers with mobile, telematics, and workflow automation.