{bc}
linkedin

Site Reliability Engineer

Tribus
Sydney, AUS
Full-time
Entry
Onsite
Discovered 1 weeks ago
Linux systems administrationPythonTerraformAnsiblePrometheusGrafana
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Linux systems administrationPythonTerraform
Smart Apply

Full Job Posting

Role Overview

Hands-on Site Reliability Engineering role supporting infrastructure for a global electronic trading platform.

Focuses on reliability, scalability, operational excellence, automation, and resilient production systems.

The team uses event-driven reliability engineering to detect abnormal behavior and infrastructure anomalies before they affect trading.

Location and Work Setting

  • The posting states Sydney or Hong Kong and onsite work in a global trading environment.

Responsibilities

  • Design, build, and maintain reliable Linux-based production infrastructure.
  • Develop infrastructure as code using Terraform and Ansible.
  • Build observability platforms with Prometheus, Grafana, and Splunk.
  • Create event-driven monitoring, intelligent alerting, and automated remediation workflows.
  • Improve incident response and operational tooling for trading systems.
  • Support PostgreSQL, InfluxDB, and Prefect production platforms.
  • Administer CEPH, VMware, KVM, and Proxmox virtualization and storage platforms.
  • Collaborate with software engineers and improve CI/CD and deployment tooling.

Requirements

  • Strong Linux systems administration and troubleshooting experience.
  • Python development experience for internal tools.
  • Experience with infrastructure as code using Terraform, Ansible, or similar tools.
  • Experience with Prometheus, Grafana, Splunk, or comparable observability platforms.
  • Experience designing operational-event monitoring, anomaly detection, and actionable alerts.
  • Networking fundamentals including TCP/IP, routing, and firewalls.
  • Experience with Docker and modern development tooling.
  • Exposure to virtualization platforms such as VMware, KVM, Proxmox, or CEPH.
  • Low-latency, distributed, or mission-critical production experience is highly regarded.

Role Highlights

  • Work on infrastructure directly supporting real-time trading.
  • Solve reliability challenges where milliseconds matter.
  • Help shape monitoring, automation, and operational engineering practices.
  • Collaborate with experienced infrastructure and software engineers.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today