Base Career helps you apply smarter for this job.
Key skills for this role
Own the production health of customer inference workloads, including availability, request success, time to first token, inter-token latency, throughput, and operational efficiency.
Establish clear service-level indicators, objectives, performance baselines, and escalation paths for production endpoints.
Build the telemetry, dashboards, alerts, and automated diagnostics needed to detect meaningful endpoint degradation before customers report it.
Create visibility across the full inference-serving path, including request queues, routing, scheduling, model servers, GPU utilization, networking, storage, and provider infrastructure.
Lead the investigation of complex latency, throughput, capacity, and reliability regressions.
Determine whether an issue originates in customer traffic patterns, platform services, inference-engine configuration, GPU hardware, networking, storage, or an external infrastructure provider.
Remain accountable for the customer outcome while partnering with the appropriate engineering teams to implement the fix.
Help lead customer-impacting incidents and establish effective operational practices for acknowledgement, diagnosis, recovery, and communication.
Convert significant incidents into automated tests, safeguards, runbooks, capacity controls, anomaly detection, and platform improvements.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
San Mateo, USA
San Francisco, USA
San Mateo, USA
San Mateo, USA
San Francisco, USA
San Mateo, USA
San Mateo, USA
San Mateo, USA
San Mateo, USA
San Francisco, USA
San Francisco, USA
Partner with the LLM Performance team to validate that engine-level optimizations deliver measurable improvements in production.
Analyze workload behavior, capacity requirements, utilization, tail latency, and cost efficiency across heterogeneous GPU providers and hardware.
Help ensure that customer performance requirements are met without consuming unnecessary infrastructure capacity.
Identify recurring patterns across incidents, workloads, and customer escalations.
Translate those findings into improvements to the inference platform, reliability architecture, deployment processes, observability, and product roadmap.
Deep experience operating critical, customer-facing or business-critical production systems.
Ability to reason about service health across multiple layers rather than treating infrastructure availability as the complete customer outcome.
Experience defining and operating service-level indicators and objectives, building actionable observability, leading incidents, performing failure analysis, and reducing mean time to detection and recovery.
Strong understanding of latency, throughput, queueing, resource contention, capacity, workload distribution, and tail-performance behavior.
Demonstrated ability to diagnose difficult production regressions and isolate bottlenecks across applications and infrastructure.
Knowledge of distributed-systems principles, including fault tolerance, scheduling, routing, load balancing, capacity management, consistency, and failure recovery.
Strong experience with Kubernetes, Linux, networking, storage, cloud infrastructure, and containerized production environments.
Experience operating across multiple cloud providers, regions, hardware configurations, or infrastructure suppliers is especially valuable.
Ability to write production-quality software and build internal tooling, instrumentation, automation, and diagnostic systems.
Proficiency in languages such as Python, Go, Java, C++, or Rust.
Experience with GPUs, ML infrastructure, model serving, vLLM, SGLang, Triton, TensorRT-LLM, or similar technologies is valuable but not required.
You should be excited to develop expertise in concepts such as time to first token, inter-token latency, continuous batching, KV caching, speculative decoding, quantization, and tokens per GPU.
Within your first six months, you will help establish clear end-to-end ownership and observability for Parasail’s production inference workloads.
Success means:
Material endpoint regressions are increasingly detected before customers report them.
Engineers can quickly determine which layer of the system is responsible for a production issue.
Customer-impacting incidents are acknowledged, diagnosed, and resolved faster.
Production endpoints consistently meet defined availability, latency, throughput, and efficiency objectives.
Recurring failure modes are converted into automated detection, safeguards, and durable platform improvements.
Infrastructure and inference capacity are used efficiently while preserving customer performance and model quality.
This role is pivotal in building a new operational discipline for AI infrastructure.
You may come from GPU infrastructure, ML platforms, databases, streaming systems, search, low-latency services, distributed storage, or large-scale production engineering. You do not need to have previously held the title “Inference Reliability Engineer.”
What matters is that you have owned important production systems, can investigate problems that span organizational and technical boundaries, and continue following an issue until the customer experience is restored.
If you are passionate about production systems, performance, reliability, and learning the cutting edge of AI infrastructure, we are excited to welcome you aboard.
AI inference cloud providing production-ready open-model services for AI-native startups.
Visit company websiteJobs and hiring trendsFull-time
Senior · 5+ years experience
Onsite
Apply faster on company sites with our extension.