{bc}
indeed

AI Engineer - Inference

Firmus Technologies
Sydney, AUS
Full-time
Onsite
Discovered 1 weeks ago
AI inferenceModel servingPythonC++ or GoTensorRT-LLM, TensorRT, SGLang, vLLM or Triton Inference ServerCUDA, cuDNN and NCCL
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

AI inferenceModel servingPython
Smart Apply

Full Job Posting

Role summary

Senior AI Engineer role focused on building and improving inference capability for internal products, external customers and future inference-as-a-service offerings.

The role establishes engineering foundations for self-hosted model serving in Firmus Technologies' AI Factory environment.

The position contributes inference workload expertise to the Model-to-Grid product and agentic applications roadmap.

Inference platform responsibilities

  • Build, operate and improve self-hosted AI inference services.
  • Define model onboarding, deployment, endpoint provisioning, testing, release and lifecycle workflows.
  • Provision secure and scalable endpoints for generation, RAG, embeddings, reranking, batch, multimodal, tool-calling and agentic use cases.
  • Develop reusable deployment templates, APIs, SDKs, configuration standards and self-service workflows.
  • Create versioned inference recipes covering models, runtimes, precision, GPU configurations, topology, scaling and benchmark results.
  • Work with Kubernetes and scheduler teams on resource profiles, placement, priorities, quotas, autoscaling and capacity policies.

Performance and operations

  • Optimize inference through quantization, compilation, batching, request routing, KV-cache management, caching, load balancing, memory optimization and distributed parallelism.
  • Build controlled benchmarking, qualification, load testing, profiling and regression testing workflows.
  • Measure latency, throughput, concurrency, GPU utilization, memory efficiency, scaling, power, cost and reliability indicators.
  • Establish observability for availability, requests, latency, errors, capacity, cost, power and service-level objectives.
  • Provide governed endpoints and performance information to agentic systems and partner teams.

Skills and experience

  • At least 5 years of software engineering experience, including at least 3 years in AI inference, model serving, ML systems, high-performance computing or distributed systems.
  • Production experience with model-serving platforms, inference APIs, GPU-backed services, AI developer platforms or multi-tenant AI systems.
  • Experience with modern inference frameworks such as TensorRT-LLM, TensorRT, SGLang, vLLM or Triton Inference Server.
  • Strong NVIDIA AI stack knowledge covering CUDA, cuDNN, NCCL, GPU profiling and distributed communication.
  • Strong Python skills and working proficiency in C++ or Go.
  • Experience with Kubernetes, containers, CI/CD, GitOps, APIs, autoscaling, observability and multi-tenant operations.
  • Understanding of GPU topology, high-performance networking, storage throughput and their impact on inference performance.
  • Understanding of inference security and governance, including identity, tenant isolation, quotas, secrets, audit logging and data protection.

Company context

Firmus Technologies develops and operates efficient AI infrastructure and a GPU cloud platform for training and deploying AI models.

The company combines AI software orchestration, energy management, liquid cooling and infrastructure operations in its AI Factory model.

Requirements

  • At least 5 years of software engineering experience, including at least 3 years in AI inference, model serving, ML systems, high-performance computing, distributed systems or comparable performance-critical environments.
  • Demonstrated experience building, operating or materially improving production model-serving platforms, inference APIs, GPU-backed services, AI developer platforms or multi-tenant AI systems.
  • Hands-on experience with one or more modern inference frameworks, including TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, Hugging Face TGI or equivalent.
  • Strong understanding of the NVIDIA AI software stack, GPU profiling, distributed communication and GPU performance analysis.
  • Practical understanding of LLM and generative-AI serving behavior, including batching, concurrency, KV-cache management, scheduling and latency-throughput trade-offs.
  • Experience with quantization, compilation, mixed precision, kernel fusion, memory optimization, caching, speculative decoding, parallelism and accuracy-performance validation.
  • Strong Python skills and working proficiency in C++ or Go.
  • Experience with distributed inference or training patterns, collective communication, fault handling and multi-node scaling.
  • Familiarity with Kubernetes, containers, CI/CD, GitOps, service APIs, autoscaling, workload scheduling, observability and production multi-tenant operations.
  • Understanding of GPU topology, NVLink, NVSwitch, PCIe, NUMA, NIC affinity, RDMA, RoCEv2, network fabrics and storage throughput.
  • Experience with inference benchmarking, profiling, load testing, regression testing and analysis of throughput, latency, utilization, scaling, power and cost metrics.
  • Understanding of security and governance for inference services, including authentication, authorization, tenant isolation, quotas, secrets, audit logging and data protection.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Firmus Technologies