Base Career helps you apply smarter for this job.
We are seeking a hands-on Principal Data & ML Ops Architect to lead the architecture of our data and document-intelligence platforms. ICA works across healthcare, federal health, regulatory and life-sciences environments, building systems that turn complex structured and unstructured data into reliable, actionable intelligence.
This role is responsible for designing scalable platforms that ingest, process, govern, and serve large volumes of data and documents for analytics, AI/ML, RAG, search, and operational applications. You will help modernize and unify existing platforms while establishing reusable architectural patterns for new solutions.
KEY RESPONSIBILITIES:
Architect large-scale data and document-processing platforms, from ingestion through downstream consumption.
Design robust patterns for document parsing, extraction, normalization, validation, storage, search, human review, and reprocessing.
Establish standards for schemas, metadata, lineage, provenance, data quality, reproducibility, and auditability.
Design reliable ingestion and transformation pipelines using APIs, files, batch processing, and streaming where appropriate.
Create platform patterns that enable Data Science and AI prototypes to become secure, maintainable production solutions.
Architect data and retrieval foundations supporting RAG, analytics, AI workflows, evaluation, and feedback loops.
Make architecture decisions involving scalability, reliability, security, privacy, cost, and operational complexity.
Lead technical design reviews, investigate production failures, mentor engineers, and work across Data Engineering, Data Science, DevOps/MLOps, Product, Security, and client teams.
Remain technically hands-on enough to review and develop Python, SQL, schemas, APIs, and production data pipelines.
REQUIRED QUALIFICATIONS:
Deep experience architecting and owning production data platforms or distributed data systems.
Strong Python, SQL, data modeling, ETL/ELT, workflow orchestration, and cloud-data-platform experience.
Significant experience with unstructured data or document-processing systems at scale.
Strong understanding of metadata, schema evolution, lineage, provenance, data quality, observability, retries, replay, backfills, and recovery.
Experience designing systems with sensitive-data controls, governance, access management, and audit requirements.
Strong understanding of how modern AI/ML and retrieval systems depend on production data architecture.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, USA
, USA
, USA
, USA
, USA
, USA
, USA
, USA
Demonstrated technical judgment, ownership, problem solving, and ability to challenge fragile or overly complex designs.
Strong technical communication and decision-making skills.
Must be authorized to work in the United States and have lived in the US for 3 or more consecutive years.
Must be able and willing to obtain a Public Trust Clearance
PREFERRED QUALIFICATIONS:
Experience in healthcare, life sciences, government, regulated industries, PHI/PII, or GDPR environments is valuable, but not required.
UNAUTHORIZED USE OF AI ASSISTIVE TECHNOLOGY: All application materials must be your own original work, and interviews must be completed independently. The use of AI-generated content or AI-assisted technologies during the application or interview process is prohibited unless interviewers provide explicit permission for specific activities during a live technical assessment. Any unauthorized use of AI will result in disqualification from the hiring process.
We are seeking a hands-on Principal Data & ML Ops Architect to lead the architecture of our data and document-intelligence platforms. ICA works across healthcare, federal health, regulatory and life-sciences environments, building systems that turn complex structured and unstructured data into reliable, actionable intelligence.
This role is responsible for designing scalable platforms that ingest, process, govern, and serve large volumes of data and documents for analytics, AI/ML, RAG, search, and operational applications. You will help modernize and unify existing platforms while establishing reusable architectural patterns for new solutions.
KEY RESPONSIBILITIES:
Architect large-scale data and document-processing platforms, from ingestion through downstream consumption.
Design robust patterns for document parsing, extraction, normalization, validation, storage, search, human review, and reprocessing.
Establish standards for schemas, metadata, lineage, provenance, data quality, reproducibility, and auditability.
Design reliable ingestion and transformation pipelines using APIs, files, batch processing, and streaming where appropriate.
Create platform patterns that enable Data Science and AI prototypes to become secure, maintainable production solutions.
Architect data and retrieval foundations supporting RAG, analytics, AI workflows, evaluation, and feedback loops.
Make architecture decisions involving scalability, reliability, security, privacy, cost, and operational complexity.
Lead technical design reviews, investigate production failures, mentor engineers, and work across Data Engineering, Data Science, DevOps/MLOps, Product, Security, and client teams.
Remain technically hands-on enough to review and develop Python, SQL, schemas, APIs, and production data pipelines.
REQUIRED QUALIFICATIONS:
Deep experience architecting and owning production data platforms or distributed data systems.
Strong Python, SQL, data modeling, ETL/ELT, workflow orchestration, and cloud-data-platform experience.
Significant experience with unstructured data or document-processing systems at scale.
Strong understanding of metadata, schema evolution, lineage, provenance, data quality, observability, retries, replay, backfills, and recovery.
Experience designing systems with sensitive-data controls, governance, access management, and audit requirements.
Strong understanding of how modern AI/ML and retrieval systems depend on production data architecture.
Demonstrated technical judgment, ownership, problem solving, and ability to challenge fragile or overly complex designs.
Strong technical communication and decision-making skills.
Must be authorized to work in the United States and have lived in the US for 3 or more consecutive years.
Must be able and willing to obtain a Public Trust Clearance
PREFERRED QUALIFICATIONS:
Experience in healthcare, life sciences, government, regulated industries, PHI/PII, or GDPR environments is valuable, but not required.
UNAUTHORIZED USE OF AI ASSISTIVE TECHNOLOGY: All application materials must be your own original work, and interviews must be completed independently. The use of AI-generated content or AI-assisted technologies during the application or interview process is prohibited unless interviewers provide explicit permission for specific activities during a live technical assessment. Any unauthorized use of AI will result in disqualification from the hiring process.
Provides AI and data analytics for federal health agencies.
Visit company website$170,000 – $174,000 / year
Full-time
Lead
Remote
Apply faster on company sites with our extension.