Base Career helps you apply smarter for this job.
Key skills for this role
The Senior AI/ML Ops Engineer is a senior individual contributor responsible for the infrastructure, pipelines, and operational reliability that power both classical machine learning (ML) and generative AI (GenAI)/agentic systems on Databricks and Snowflake. This role blends platform and pipeline engineering, classical ML operations, and GenAI/agentic systems infrastructure, with hands-on depth in artificial intelligence (AI) and/or ML.
You will own the reliability of ML and GenAI systems in production, from ML/AI-specific pipeline and deployment workflows through model lifecycle management to the infrastructure behind retrieval and agentic tooling, taking AI engineering and data science work from prototype to reliable production systems. This role serves as the primary operational support for AI Engineering and agentic work, and coordinates with platform and infrastructure teams on things like environment provisioning, access configuration, and underlying CI/CD tooling.
Your day to day
Platform & Pipeline Engineering
Deploy and promote ML and GenAI models, pipelines, and code across environments, using CI/CD infrastructure maintained by Infra/DevOps
Develop and promote reusable deployment patterns and tooling to reduce the effort required to stand up new AI/ML use cases and client-specific deployments
Build and maintain data pipelines supporting both classical ML and GenAI workloads, from ingestion through feature engineering to serving
Operate within Databricks and Snowflake governance frameworks (e.g. Unity Catalog access controls, environment boundaries) to ensure secure, compliant promotion of code, data, and models, and support governance practices/tooling more broadly as the team's AI/ML footprint grows
Independently diagnose and resolve production issues across pipelines, infrastructure, and model-serving systems
Classical ML Operations
Automate and monitor production ML inference and feature engineering workflows, including alerting and incident response
Own model lifecycle management using a model registry tool such as MLflow, along with Unity Catalog: experiment tracking, model registration, versioning, and controlled promotion across environments
GenAI & Agentic Infrastructure
Build and maintain infrastructure for retrieval-augmented generation (RAG) systems, including vector search indexing and retrieval pipelines
Deploy, host, and maintain MCP servers and tool integrations for agentic applications
Build and maintain evaluation infrastructure for AI systems, and contribute to evaluation methodology in partnership with AI engineering
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Pune, IND
, USA
Pune, IND
Pune, IND
, USA
, USA
, USA
, USA
Support agent observability: logging, tracing, and monitoring for agent and model behavior in production
Cross-Functional
Partner with business stakeholders to scope data and feature requirements
Coordinate with Software Engineering, Data Engineering, Data Science, security, and DevOps on infrastructure changes and shared platform needs
Mentor junior engineers on platform practices and operational standards
Occasionally contribute to customer-specific implementation work as part of a broader team (e.g. semantic layer configuration, domain-specific analytics builds such as Rx trend reporting)
What you bring to the team
5+ years of experience in AI/ML engineering, MLOps, or a closely related discipline
Extensive hands-on experience in AI/ML, with meaningful depth in at least one of the following, and some working exposure to the other:
AI/GenAI: agentic frameworks (e.g. LangChain), RAG systems, vector search, MCP or comparable tool-integration protocols, model serving/gateway layers, evaluation design
Classical ML: model development across common algorithm families (e.g. XGBoost, gradient boosting, random forest, neural networks), feature engineering, training pipelines, production deployment
Deep, hands-on experience with Databricks and/or Snowflake, including ML/AI pipeline development, working within governance/access-control frameworks (e.g. Unity Catalog), and integrating with existing CI/CD infrastructure
Strong hands-on experience with a model lifecycle/registry tool such as MLflow: experiment tracking, model registration, versioning, and promotion across environments
Strong proficiency in Python and SQL
Demonstrated ability to independently diagnose and resolve production infrastructure issues
Excellent communication skills and comfort working directly with non-technical stakeholders
A track record of proactively learning new tools and frameworks and applying them quickly to real work
What we would like to see, but not required
Experience in healthcare, health insurance, or regulated data environments
Experience building or operating multi-agent systems
Experience with Mosaic AI Gateway or comparable model-serving/gateway platforms
Compensation
Compensation for this role is based on experience, skills, and location, and includes base salary plus eligibility for performance bonuses and equity grants.
What you'll get in return
Unlimited paid time off — recharge when you need it.
Work from anywhere — flexibility to fit your life.
Comprehensive health coverage — multiple plan options to choose from.
Equity for every employee — share in our success.
Growth-focused environment — your development matters here.
Home office setup allowance — one-time support to get you started.
Monthly cell phone allowance — stay connected with ease.
Our Commitment as an Equal Opportunity Employer
As a mission-led technology company helping to drive better healthcare outcomes, Abacus Insights believes that the best innovation and value we can bring to our customers comes from diverse ideas, thoughts, experiences, and perspectives. Therefore, we dedicate resources to building diverse teams and providing equal employment opportunities to all applicants. Abacus prohibits discrimination and harassment regarding race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.
At the heart of who we are is a commitment to continuously and intentionally building an inclusive culture—one that empowers every team member across the globe to do their best work and bring their authentic selves. We carry that same commitment into our hiring process, aiming to create an interview experience where you feel comfortable and confident showcasing your strengths. If there’s anything we can do to support that—big or small—please let us know.
AI Use in Recruitment
We use AI-powered tools throughout our recruitment process. This includes Greenhouse's Real Talent Matching Technology, which helps prioritize applications based on job-related criteria, as well as additional AI tools used to support sourcing, interview preparation, and other recruiting tasks. We do not remove the human element—our recruiters and hiring teams review all applications and make all decisions related to candidate progression and hiring, regardless of which tools are used to support that process.
By applying, you acknowledge that we will collect and process your personal information for recruiting purposes in accordance with applicable data protection laws. Please review our Applicant Privacy Notice for more information, including your rights.
The Senior AI/ML Ops Engineer is a senior individual contributor responsible for the infrastructure, pipelines, and operational reliability that power both classical machine learning (ML) and generative AI (GenAI)/agentic systems on Databricks and Snowflake. This role blends platform and pipeline engineering, classical ML operations, and GenAI/agentic systems infrastructure, with hands-on depth in artificial intelligence (AI) and/or ML.
You will own the reliability of ML and GenAI systems in production, from ML/AI-specific pipeline and deployment workflows through model lifecycle management to the infrastructure behind retrieval and agentic tooling, taking AI engineering and data science work from prototype to reliable production systems. This role serves as the primary operational support for AI Engineering and agentic work, and coordinates with platform and infrastructure teams on things like environment provisioning, access configuration, and underlying CI/CD tooling.
Your day to day
Platform & Pipeline Engineering
Deploy and promote ML and GenAI models, pipelines, and code across environments, using CI/CD infrastructure maintained by Infra/DevOps
Develop and promote reusable deployment patterns and tooling to reduce the effort required to stand up new AI/ML use cases and client-specific deployments
Build and maintain data pipelines supporting both classical ML and GenAI workloads, from ingestion through feature engineering to serving
Operate within Databricks and Snowflake governance frameworks (e.g. Unity Catalog access controls, environment boundaries) to ensure secure, compliant promotion of code, data, and models, and support governance practices/tooling more broadly as the team's AI/ML footprint grows
Independently diagnose and resolve production issues across pipelines, infrastructure, and model-serving systems
Classical ML Operations
Automate and monitor production ML inference and feature engineering workflows, including alerting and incident response
Own model lifecycle management using a model registry tool such as MLflow, along with Unity Catalog: experiment tracking, model registration, versioning, and controlled promotion across environments
GenAI & Agentic Infrastructure
Build and maintain infrastructure for retrieval-augmented generation (RAG) systems, including vector search indexing and retrieval pipelines
Deploy, host, and maintain MCP servers and tool integrations for agentic applications
Build and maintain evaluation infrastructure for AI systems, and contribute to evaluation methodology in partnership with AI engineering
Support agent observability: logging, tracing, and monitoring for agent and model behavior in production
Cross-Functional
Partner with business stakeholders to scope data and feature requirements
Coordinate with Software Engineering, Data Engineering, Data Science, security, and DevOps on infrastructure changes and shared platform needs
Mentor junior engineers on platform practices and operational standards
Occasionally contribute to customer-specific implementation work as part of a broader team (e.g. semantic layer configuration, domain-specific analytics builds such as Rx trend reporting)
What you bring to the team
5+ years of experience in AI/ML engineering, MLOps, or a closely related discipline
Extensive hands-on experience in AI/ML, with meaningful depth in at least one of the following, and some working exposure to the other:
AI/GenAI: agentic frameworks (e.g. LangChain), RAG systems, vector search, MCP or comparable tool-integration protocols, model serving/gateway layers, evaluation design
Classical ML: model development across common algorithm families (e.g. XGBoost, gradient boosting, random forest, neural networks), feature engineering, training pipelines, production deployment
Deep, hands-on experience with Databricks and/or Snowflake, including ML/AI pipeline development, working within governance/access-control frameworks (e.g. Unity Catalog), and integrating with existing CI/CD infrastructure
Strong hands-on experience with a model lifecycle/registry tool such as MLflow: experiment tracking, model registration, versioning, and promotion across environments
Strong proficiency in Python and SQL
Demonstrated ability to independently diagnose and resolve production infrastructure issues
Excellent communication skills and comfort working directly with non-technical stakeholders
A track record of proactively learning new tools and frameworks and applying them quickly to real work
What we would like to see, but not required
Experience in healthcare, health insurance, or regulated data environments
Experience building or operating multi-agent systems
Experience with Mosaic AI Gateway or comparable model-serving/gateway platforms
Compensation
Compensation for this role is based on experience, skills, and location, and includes base salary plus eligibility for performance bonuses and equity grants.
What you'll get in return
Unlimited paid time off — recharge when you need it.
Work from anywhere — flexibility to fit your life.
Comprehensive health coverage — multiple plan options to choose from.
Equity for every employee — share in our success.
Growth-focused environment — your development matters here.
Home office setup allowance — one-time support to get you started.
Monthly cell phone allowance — stay connected with ease.
Our Commitment as an Equal Opportunity Employer
As a mission-led technology company helping to drive better healthcare outcomes, Abacus Insights believes that the best innovation and value we can bring to our customers comes from diverse ideas, thoughts, experiences, and perspectives. Therefore, we dedicate resources to building diverse teams and providing equal employment opportunities to all applicants. Abacus prohibits discrimination and harassment regarding race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.
At the heart of who we are is a commitment to continuously and intentionally building an inclusive culture—one that empowers every team member across the globe to do their best work and bring their authentic selves. We carry that same commitment into our hiring process, aiming to create an interview experience where you feel comfortable and confident showcasing your strengths. If there’s anything we can do to support that—big or small—please let us know.
AI Use in Recruitment
We use AI-powered tools throughout our recruitment process. This includes Greenhouse's Real Talent Matching Technology, which helps prioritize applications based on job-related criteria, as well as additional AI tools used to support sourcing, interview preparation, and other recruiting tasks. We do not remove the human element—our recruiters and hiring teams review all applications and make all decisions related to candidate progression and hiring, regardless of which tools are used to support that process.
By applying, you acknowledge that we will collect and process your personal information for recruiting purposes in accordance with applicable data protection laws. Please review our Applicant Privacy Notice for more information, including your rights.
Healthcare data management platform that simplifies data integration, improves data quality, and delivers actionable insights for health plans and providers.
Visit company websiteJobs and hiring trendsSenior · 5+ years experience
Apply faster on company sites with our extension.