Senior Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Who We Are NEXT Ventures is a global fintech group powering FundedNext — one of the world's fastest-growing proprietary trading platforms — and FNmarkets, a regulated CFD brokerage.
Across offices in Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai, we build and operate the technology that lets traders access global markets at scale.
Our Platform Engineering team owns the infrastructure, reliability, and observability backbone that every product squad depends on.
Key Skills for This Role
Full Job Posting
Job description
Who We Are
NEXT Ventures is a global fintech group powering FundedNext — one of the world's fastest-growing proprietary trading platforms — and FNmarkets, a regulated CFD brokerage. Across offices in Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai, we build and operate the technology that lets traders access global markets at scale. Our Platform Engineering team owns the infrastructure, reliability, and observability backbone that every product squad depends on.
Your Role in Our Mission
As our Site Reliability Engineer, you are the dedicated specialist who keeps our services observable, fast, and resilient. You own the centralized logging and alerting backbone, drive service-level optimization across the stack, and perform log analysis across both Linux and Windows environments. Working within the Platform Engineering squad, you execute the reliability initiatives that free the squad lead to focus on architecture — and you are the reason incidents are short, signals are clean, and detection is fast.
This role is distinct from our DevSecOps Engineer: DevSecOps builds and secures the platform foundation; you measure, detect, and optimize on top of it.
How You'll Make an Impact
Centralized Logging & Alerting
Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes.
Drive MTTD under 15 minutes through better signals, dashboards, and runbooks.
Service-Level Optimization & Reliability
Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers — profile before guessing, measure every fix.
Define, implement, and own SLOs, SLIs, and error budgets for critical customer-facing services, and drive improvements against them.
Build deep observability with Datadog — APM, dashboards, monitors, log management, and SLO tracking.
Lead reliability and performance root-cause analysis and drive durable fixes.
Support load testing and capacity planning — identify breaking points before traffic growth causes production issues.
Operations & Cross-Team Support
Participate in a shared on-call rotation with solid runbooks and blameless post-incident reviews.
Continuously reduce manual toil through automation and better tooling.
Collaborate with product squads to instrument services, define meaningful SLIs, and surface the right signals.
Document runbooks, dashboards, and operational procedures so any engineer can respond to incidents with clear guidance.
What You Bring
5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or strongly reliability-focused DevOps work.
Strong hands-on experience with centralized log management platforms — ingestion, parsing, structured logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.
Able to diagnose incidents through log analysis across both Linux and Windows environments, isolating root causes under pressure.
Designs alerting systems with tiered thresholds that minimize noise while catching real problems early, with clear escalation paths.
Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.
Experienced defining and operating SLOs, SLIs, and error budgets for customer-facing services.
Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix.
Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.
Hands-on experience operating services on Kubernetes, EKS preferred.
Familiar with Infrastructure-as-Code tooling such as Terraform for collaboration with the DevSecOps team.
Evidence-driven — profiles and measures before guessing; validates every optimization against before/after data.
Reliability-oriented — treats detection speed and signal quality as first-class engineering problems.
Good communicator — works with product squads to define SLIs and explains reliability constraints clearly.
X-Factor: AI-Native Engineering
You actively use modern AI agentic workflows daily — not limited to Copilot autocomplete.
You are proficient with Claude Code, Cursor, Windsurf, or equivalent tools.
You are comfortable with project-level AI configuration ( CLAUDE.md , rules files), agentic task delegation, and AI-driven code review.
You think in terms of 5–10x productivity through AI-augmented development — and you can demonstrate it.
Your Journey After Applying
Stage 1 — TA Interview
Stage 2 — Screening Questionnaire
Stage 3 — Hiring Manager Interview
Stage 4 — Head of IT Interview
Why Join NEXT
Work on reliability challenges at real scale — 100M+ row data stores, multi-region traffic, high-frequency trading infrastructure.
A team that treats observability and detection speed as first-class engineering problems, not afterthoughts.
Flat structure — your work directly shapes how the platform operates, not filtered through layers of process.
Offices across Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai — a genuinely global engineering team.
Competitive compensation benchmarked to your market, with room to grow as the team scales.
About NEXT Ventures
UAE-based fintech group operating a prop-trading platform and CFD brokerage for traders worldwide.
Visit company websiteJobs and hiring trendsApply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at NEXT Ventures
Senior Risk Manager
Dubai, UAE
NEXT Ventures is seeking a Senior Risk Manager to lead trading risk, investigations, surveillance, and analytics across CFD, futures, and crypto products. The role requires 7+ years of relevant experience, leadership of
CRM Lead
, UAE
NEXT Ventures is seeking a CRM Lead to build its CRM foundation, own the customer lifecycle, and improve trader activation, retention, and re-engagement. The role requires hands-on experience building CRM programs, end-t
Senior Personal Assistant
, UAE
NEXT Ventures is seeking a Senior Personal Assistant to support the CEO and CSO across executive, personal, family, travel, and logistical matters. The role requires 5+ years supporting senior executives, exceptional dis
Assistant Manager — Employee Experience & Administration
Dubai, UAE
Trading Platform Engineer
Sydney, AUS
Senior Risk Manager
Dubai, UAE
CRM Lead
Dubai, UAE
CRM Lead
, UAE
Senior Personal Assistant
, UAE
Trading and Risk Advisor — CFD
Dubai, UAE
Assistant Manager — Employee Experience & Administration
Dubai, UAE
Trading and Risk Advisor - Futures
, IND