Base Career helps you apply smarter for this job.
Key skills for this role
As a Monitoring and Observability Engineer, you'll help keep OQC's live quantum computing systems running at their best. By improving system visibility, developing intelligent monitoring solutions and driving operational improvements, you'll play a key role in maximising uptime and enabling world-class live quantum computing services.
Working within our Live Services team, you'll monitor the health and performance of our live cryogenic systems, respond to operational incidents and continuously improve our observability capabilities. You'll collaborate across engineering, operations, software and reliability teams to develop dashboards, refine alerting strategies and automate operational responses that improve reliability and reduce downtime.
Monitor live systems, triage operational alerts and provide first-line incident response to maximise system availability.
Develop and maintain dashboards that deliver clear visibility into system health, trends and operational performance.
Analyse recurring incidents and monitoring data to identify root causes, improve alert quality and drive preventative action.
Collaborate with Reliability Engineering, Software and Operations teams to enhance observability tooling and operational processes.
Define monitoring thresholds, alerting strategies and operational metrics that support reliable live quantum computing services.
Produce documentation covering monitoring processes, incident response, escalation procedures and operational handovers.
Contribute to automation and continuous improvement initiatives that reduce manual intervention and improve service reliability.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Reading, GBR
London, GBR
Reading, GBR
Reading, GBR
London, GBR
Reading, GBR
Reading, GBR
Reading, GBR
Reading, GBR
Experience working in monitoring, observability or operational support for complex technical systems.
Experience working with dashboards, telemetry, alarms, logs or time-series data.
Experience supporting uptime, incident response or on-call operational environments.
Strong analytical and problem-solving skills with the ability to identify trends, anomalies and operational risks.
Experience developing software in Python using modern development practices.
Experience creating and maintaining operational dashboards.
Excellent communication skills with the ability to collaborate effectively across technical teams.
Degree, HNC/HND, apprenticeship or equivalent experience in Engineering, Physics, Computer Science, Controls, Instrumentation, Data or a related technical discipline.
Willingness to participate in an on-call rota and travel internationally when required.
Experience with Grafana Labs or similar observability platforms.
Experience monitoring high-availability technical systems.
Understanding of incident management, root cause analysis, reliability engineering or Site Reliability Engineering (SRE) principles.
Experience with HTTP/REST APIs, Git, containerisation, Kubernetes or Infrastructure as Code.
Continuous improvement, ownership or leadership experience.
You will join a world-class team at the forefront of the next computational era. We offer a culture of bold innovation, the chance to work with unique lab infrastructure, and the opportunity to see your work redefine the limits of computation.
Learn more about our benefits and positive work culture here: https://oqc.tech/company/careers-at-oqc/
UK-based quantum computing company operating superconducting quantum computers in data centres for enterprise and government customers.
Visit company websiteJobs and hiring trendsFull-time
Mid
Hybrid
Apply faster on company sites with our extension.