Base Career helps you apply smarter for this job.
Key skills for this role
What You’ll Do:
As part of the BQL (Bloomberg Query Language) Reliability Engineering team, you will build software and platform capabilities that improve the reliability, resilience, and transparency of BQL and the services it depends on. You’ll work on engineering problems at significant scale, using software and automation to make reliability a built-in property of the platform rather than a purely operational concern.
You’ll be trusted to:
Design, build, and maintain software and self-service platform capabilities that enable engineering teams to understand, operate, and improve the reliability of BQL at scale.
Build tools and automated diagnostic capabilities that analyze telemetry and system behavior, helping engineers rapidly identify failures, regressions, and their root causes.
Develop software that improves incident detection and diagnosis, reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) for high-severity incidents.
Engineer observability capabilities that turn metrics, logs, traces, and other system signals into actionable insights across BQL’s distributed architecture.
Partner with BQL engineering teams on system design, instrumentation, SLIs and SLOs, ensuring reliability is built into services throughout the Software Development Lifecycle.
Improve platform resilience through engineering and experimentation, including load and stress testing, canary releases, controlled experiments, and failure testing.
Identify recurring operational problems and eliminate them through software, automation, and improvements to platform architecture.
You’ll Need to Have:
4+ years of experience in Software Engineering, Reliability Engineering, Platform Engineering, or a related technical role.
Experience working with an object-oriented programming language (C/C++, Python, Java, etc.)
Strong knowledge of Linux/UNIX systems and experience developing or operating distributed applications in production.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
London, GBR
New York City, USA
New York City, USA
London, GBR
London, GBR
London, GBR
London, GBR
New York City, USA
New York City, USA
New York City, USA
New York City, USA
Demonstrated experience improving the performance, availability, resilience, or scalability of mid- to large-scale systems.
Experience with production software delivery, including deployment, release management, testing, and safely introducing changes into distributed systems.
Strong analytical and problem-solving skills, including the ability to use production data and system telemetry to understand complex system behavior.
BA, BS, MS, or PhD in Computer Science, Engineering, or a related technical field.
We’d Love to See Experience with:
Building developer platforms, reliability tooling, observability systems, or other infrastructure used by engineering teams.
Working with metrics, logs, distributed traces, time-series data, and other forms of production telemetry.
Defining and applying SLIs, SLOs, and other quantitative measures of system reliability.
Designing and executing load tests, stress tests, failure tests, canary releases, or other techniques for validating system resilience.
Applying statistical methods to understand system behavior and solve real-world engineering problems.
Querying and analyzing large-scale datasets in enterprise data environments.
Designing, executing, and analyzing A/B tests and other controlled experiments.
Operating in regulated or highly controlled environments.
Bloomberg is a global financial data, technology, and media company providing analytics and tools for financial professionals, including the Bloomberg Terminal.
Visit company websiteJobs and hiring trendsUSD 160000-240000 / year
Senior · 4+ years experience
Apply faster on company sites with our extension.