Base Career helps you apply smarter for this job.
Key skills for this role
Baselayer is building the most comprehensive, accurate, and continuously-current identity graph of US businesses — fusing public records, IRS data, sanctions lists, web signals, and fraud telemetry from thousands of financial institutions into a single graph that resolves any business in milliseconds. None of that works without world-class data infrastructure. We’re hiring a Data Engineer to help build and run the pipelines and models that turn messy, heterogeneous data into trustworthy, production-grade signal. You’ll write real production code in your first weeks, own pipelines end to end, and learn alongside senior data and ML engineers who will invest in your growth. This is a role for an early-career engineer who wants to be close to the action: feeding the models, not just cleaning up after them.
Every business in America needs a bank account to exist. The system that decides whether they're real, who's behind them, and whether they're a risk, runs on infrastructure from the 1980s. We're rebuilding that layer from scratch.
Baselayer is the identity layer for institutions across the United States — the most complete business graph in America and every human tied to it. We fuse public records, IRS data, sanctions lists, web signals, and fraud telemetry from 2,200+ financial institutions into a single graph that resolves any business and the humans behind it in milliseconds. The legacy credit bureaus took 50 years to build something that gets 60% match rates. We've built something that gets 98% in under two years.
Today we're trusted by over 20% of financial institutions in America — including FIS, Rho, Socure and leading loan infrastructure providers. But the graph is becoming infrastructure for anyone who needs to know if a business is real and worth trusting: gig platforms, marketplaces, AI companies, and commerce infrastructure at scale.
Trust is the substrate of every financial transaction. We're rebuilding it.
We're solving real-time entity resolution at a scale no one else has cracked — fusing dozens of data sources into a single business identity graph and resolving any entity in milliseconds. It's a graph AI problem, a retrieval problem, and a fraud-modeling problem stacked on top of each other. The technical depth is real.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
San Francisco, USA
New York City, USA
San Francisco, USA
New York City, USA
New York City, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
You'd be joining a small team where the data moat is defensible, the research problems are open, and the infrastructure you build becomes load-bearing for businesses. Ownership is real. Velocity is real. There's no layer of process between an idea and shipping it.
We're at an inflection point — the graph is built, the match rates speak for themselves, and the hardest problems are still ahead: graph embeddings, fraud propagation models across the business network, real-time traversal at sub-100ms latency, and expanding the identity layer beyond finance into every platform that needs to trust a business.
If you want to work on something foundational — the kind of infrastructure that gets built once and everything else runs on top of — this is it.
Baselayer is building the most comprehensive, accurate, and continuously-current identity graph of US businesses — fusing public records, IRS data, sanctions lists, web signals, and fraud telemetry from thousands of financial institutions into a single graph that resolves any business in milliseconds. None of that works without world-class data infrastructure. We’re hiring a Data Engineer to help build and run the pipelines and models that turn messy, heterogeneous data into trustworthy, production-grade signal. You’ll write real production code in your first weeks, own pipelines end to end, and learn alongside senior data and ML engineers who will invest in your growth. This is a role for an early-career engineer who wants to be close to the action: feeding the models, not just cleaning up after them.
Build and maintain ETL/ELT pipelines that ingest and normalize public records, web signals, and fraud telemetry from dozens of sources
Develop data models and transformation layers (Dataflow, Spark, Airflow) that power fraud detection, KYB, and customer-facing APIs
Implement data quality checks, observability tooling, and alerting so problems surface before customers see them
Tune pipelines and queries for performance, freshness, and cost in our cloud data warehouse
Work with data scientists, ML engineers, and product to make clean, well-modeled data available for entity resolution and scoring
Help ensure pipelines meet security and regulatory standards for sensitive data (SOC 2, GDPR, KYC/KYB)
Document what you build and translate between technical and non-technical stakeholders so the rest of the team moves faster
Curiosity about AI/ML infrastructure and a desire to be close to the models, not just the cleanup after them
Experience with streaming or real-time data systems (e.g. Kafka, Pub/Sub)
Exposure to KYC/KYB, fraud, risk, or underwriting data, and the ethical care that sensitive information demands
GCP experience (BigQuery, Cloud Run, Dataflow, Pub/Sub)
You care deeply about data quality and trust, and build systems others can rely on
You’ve worked without a playbook before, and you take direct feedback well and act on it fast
AI-powered business identity and risk platform serving banks, lenders, fintechs, and other financial institutions.
Visit company websiteJobs and hiring trendsUSD 120000-150000 yearly / year
Full-time
Entry · 1+ years experience
Hybrid
Apply faster on company sites with our extension.