{bc}
lever

AI Platform Engineer

jobgether
Remote, IND
Full-time
Remote
Discovered 1 weeks ago
PythonGo or JavaDistributed systemsAI platform engineeringMLOps and LLMOpsLarge language model systems
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

PythonGo or JavaDistributed systems
Smart Apply

Full Job Posting

Role overview

Build a central AI platform from the ground up for multiple product teams.

Develop infrastructure, services, and tooling for reliable AI delivery at scale.

Work across AI platform engineering, RAG, LLMOps, observability, security, automation, and production operations.

Make hands-on technical decisions shaping how AI systems are built, deployed, evaluated, and operated.

Accountabilities

  • Design, build, deploy, and operate scalable AI platform services with API contracts, versioning, SLOs, and production-readiness standards.
  • Develop production-grade Python services, shared libraries, SDKs, and integration patterns.
  • Build a multi-provider model gateway with routing, fallback, retry, rate limiting, cost attribution, and budget enforcement.
  • Develop RAG-as-a-service covering ingestion, chunking, embeddings, retrieval, hybrid search, and quality measurement.
  • Establish prompt management, guardrails, content safety, and reliable agent and tool-use patterns.
  • Build observability for tracing, latency, errors, cost, quality, drift, and audit logging.
  • Maintain evaluation infrastructure with benchmarks, LLM-as-judge methods, human review, and regression testing.
  • Operate model deployment pipelines with safe release and rollback processes.
  • Establish alerting and SLO frameworks, participate in on-call rotations, and improve incident response.
  • Partner with product engineering teams on onboarding, integration guidance, troubleshooting, architecture, vendor evaluation, and platform improvement.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • At least 5 years of professional software engineering experience with a strong production track record.
  • Approximately 6 or more years building production distributed systems, ideally platforms or API gateways at scale.
  • At least 5 years in MLOps, LLMOps, ML platform engineering, or comparable production DevOps and machine-learning engineering.
  • At least 2 years of hands-on production experience with LLM-based systems.
  • Strong Python plus Go or Java development skills for production-quality code.
  • Strong understanding of asynchronous programming, backpressure, rate limiting, distributed systems, and service design.
  • Experience with LLM observability, evaluation methodologies, multi-tenant systems, and isolation guarantees.
  • Cloud-native experience with AWS or Azure, Kubernetes, service mesh, Terraform, infrastructure-as-code, and CI/CD.
  • Experience with safe model deployment techniques and understanding of the full ML lifecycle.
  • Strong AI security knowledge covering PII, tenant isolation, prompt injection, and audit logging.
  • Fluent English communication skills.

Benefits and work environment

  • Fully remote opportunity available across India.
  • Work schedule aligned to 11:00 AM–8:00 PM.
  • Join a newly formed AI platform team with significant technical ownership.
  • Gain exposure to modern LLM, RAG, AI agent, MLOps, LLMOps, cloud, observability, and platform technologies.
  • Work on production AI systems used across multiple product lines in a distributed engineering environment.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at jobgether