Base Career helps you apply smarter for this job.
Key skills for this role
● Own the health of Sigma, MiQ's enterprise platform, by building world-class observability across its services and infrastructure.
● You will design and implement monitoring, alerting, and synthetic checks that proactively surface application errors, performance regressions, and infrastructure issues before our users notice them — and partner with product engineering teams to drive them to resolution.
● You'll work hands-on with our Grafana-based observability stack (metrics, logs, traces, and synthetic monitoring) and our AWS/Kubernetes (EKS) environment to define meaningful SLIs and SLOs, reduce alert noise, build actionable dashboards and runbooks, and strengthen incident response practices including on-call readiness.
● You will improve release safety and platform stability by designing and optimizing release pipelines, enabling progressive delivery, and implementing health gates, automated deployment analysis, and feature-flag rollbacks to detect issues early and minimize user impact.
● You will bring a performance engineering mindset — load testing, capacity analysis, latency and resource profiling — and automate away toil through scripting and infrastructure-as-code.
● You will accelerate SRE maturity by applying AIOps capabilities to improve signal quality, speed up detection and diagnosis, and reduce manual operational effort. ● Over time, you will grow into owning the observability charter for the platform, setting standards and evangelizing best practices across engineering teams.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Brisbane, AUS
Toronto, CAN
Toronto, CAN
Toronto, CAN
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Key stakeholders include the DevOps & Cloud team, Sigma product engineering teams, Tech Leads across Engineering, and engineering leadership who rely on reliability and performance insights to make informed decisions.
• 4 - 8 years of experience in Site Reliability Engineering, DevOps, or platform/production engineering roles supporting customer-facing systems.
• Hands-on experience with observability tooling — Grafana, Prometheus, Datadog, and log/trace aggregation (e.g., Loki, Tempo, OpenTelemetry, or equivalent) — covering metrics, logs, traces, and events.
• Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively.
• Deep knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews.
• Solid experience operating workloads on Kubernetes (ideally EKS) and AWS — comfortable debugging issues across the application, container, and infrastructure layers.
• Strong scripting and automation skills in Python, Bash, or Go, with exposure to infrastructure-as-code (e.g., Terraform) and CI/CD pipelines.
• A performance engineering orientation: load/stress testing (e.g., k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks.
• Experience contributing to incident management — triage, escalation, communication, and post-mortems — and helping set up or improve on-call processes and rotations.
• A proactive, ownership-driven mindset: you identify problems from telemetry before they're reported, and you follow through with the teams who need to fix them.
• Clear written and verbal communication — you can turn noisy signals into crisp findings, runbooks, and recommendations for engineering teams.
• Strong communication and collaboration abilities to work effectively alongside peer Platform teams, such as Cloud, DevOps, and Security.
We’ve highlighted some key skills, experience and requirements for this role. But please don’t worry if you don’t meet every single one. Our talent team strives to find the best people. They might see something in your background that fits this role, or another opportunity at MiQ.
You will be the engineer who makes reliability visible and actionable for MiQ's flagship platform. You'll build the observability foundations — well-thought-out dashboards, meaningful alerts, synthetic checks, and SLOs — that let teams detect and resolve errors, performance degradations, and infrastructure issues before they impact users. You'll reduce mean time to detect and resolve, cut alert fatigue, and raise the bar on incident response and on-call readiness. As you grow, you'll own the observability charter: defining standards, introducing new tooling and practices where appropriate, and evangelizing a proactive reliability culture across engineering.
Our Center of Excellence is the very heart of MiQ, and it’s where the magic happens. It means everything you do and create will have a huge impact on our entire global business.
MiQ is incredibly proud to foster a welcoming culture. We do everything possible to make sure everyone feels valued for what they bring. With global teams committed to diversity, equity, and inclusion, we’re always moving towards becoming an even better place to work.
Our values are so much more than statements . They unite MiQers in every corner of the world. They shape the way we work and the decisions we make. And they inspire us to stay true to ourselves and to aim for better. Our values are there to be embraced by everyone so that we naturally live and breathe them. Just like inclusivity , our values flow through everything we do - no matter how big or small.
• Generous annual PTO paid parental leave, with two additional paid days to acknowledge holidays, cultural events, or inclusion initiatives. • Employee resource groups are designed to connect people across all MiQ regions, drive action, and support our communities.
Verified company details for this employer are not available yet.
Full-time
Mid · 4+ years experience
Hybrid
Apply faster on company sites with our extension.