{bc}
company_site

Software Engineer III, Site Reliability Engineering - Ventures

Electronic Arts
Guildford, GBR
Mid
Hybrid
Discovered 4 days ago
awsbashcloudwatchdatadogdockergcp
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

awsbashcloudwatch
Smart Apply

Full Job Posting

Key Responsibilities:

Observability and incident response (primary focus)

• Design and build observability across Ventures environments, including logging, metrics, tracing, dashboards, and alerting.

• Define meaningful service level indicators and alert thresholds so that teams are paged on real problems rather than noise.

• Lead and improve incident response practice, including triage, escalation paths, on-call rotation, and runbooks.

• Run post-incident reviews and drive the follow-up actions that come out of them.

• Give product engineering teams the visibility they need to diagnose issues in their own services.

Cloud cost and governance

• Establish visibility of cloud spend across Ventures projects in AWS and GCP.

• Identify and implement cost efficiency improvements, and set up guardrails to prevent regressions.

• Define and enforce tagging, account, and resource conventions so that spend and ownership are attributable.

• Report on cost trends and forecast implications for engineering leadership.

Security posture

• Review and improve secrets management, IAM and least-privilege access, and network configuration.

• Work with EA security and platform teams to align Ventures infrastructure with internal standards.

• Surface and help remediate infrastructure security risks across projects.

Agentic operations

• Build and refine agentic workflows that automatically inspect incidents, correlate telemetry, and propose remediations for human confirmation.

• Make our systems legible to agents, including structured telemetry, machine-readable runbooks, and well-scoped tooling.

• Review agent-proposed actions and remain accountable for what gets applied, treating suggestions as input rather than answers.

• Track where agentic workflows perform well and where they do not, and feed that back into how they are designed.

Infrastructure and delivery support

• Support the backend and platform team on Terraform and infrastructure as code, contributing modules and improvements where useful.

• Support and improve existing CI/CD pipelines where they create friction for product teams, without taking over ownership.

• Integrate with internally provided EA infrastructure and platform services, and work with those teams to unblock delivery.

• Document your work, including runbooks and operational procedures, to a standard the team can own after the contract ends.

How Your Success Will Be Measured:

• Coverage and usefulness of monitoring and alerting across Ventures environments.

• Time to detect, respond to, and resolve incidents, and the quality of post-incident follow-through.

• Demonstrable reduction or better control of cloud spend, with clear attribution by project.

• Measurable improvements in security posture across Ventures infrastructure.

• Reduced operational friction for product engineering teams.

• Growth in the coverage and accuracy of agentic workflows, and in the confidence the team has in acting on their output.

• Quality of documentation and handover, measured by the team’s ability to operate independently at contract end.

• Effective collaboration with the backend and platform team and with EA platform and security teams.

What You Bring:

  • 5+ years of professional SRE, DevOps, platform, or infrastructure engineering experience.
  • Strong experience implementing observability using tools such as Datadog, Grafana, Prometheus, OpenTelemetry, or CloudWatch.
  • Proven experience running incident response in a production environment, including on-call and post-incident review.
  • Experience managing and optimising cloud cost at organisational scale, including tagging and governance practices.
  • Solid understanding of cloud security fundamentals, including secrets management, least-privilege access, and network segmentation.
  • Strong hands-on AWS experience across compute, networking, storage, IAM, and managed services.
  • Working knowledge of GCP, or clear evidence of operating comfortably in a multi-cloud environment.
  • Practical Terraform experience, sufficient to contribute confidently to an existing infrastructure as code estate.
  • Familiarity with CI/CD tooling such as GitHub Actions, GitLab CI, Jenkins, or similar.
  • Experience with containerisation and orchestration, including Docker and Kubernetes.
  • Strong scripting ability in Python, Bash, Go, or similar.
  • Experience working alongside internal platform or shared infrastructure teams within a larger organisation.
  • Hands-on experience using AI coding or operations agents in real work, and a clear view on where they help and where they should not be trusted.
  • A track record of joining an existing environment and becoming productive quickly.

Soft Skills & Values:

• Self-directed and comfortable making progress with limited ramp-up time.

• Pragmatic, with good judgement about what to fix now and what to document and leave.

• Genuinely enthusiastic about agentic workflows, and equally willing to say when an agent has got it wrong.

• Comfortable improving systems owned by other teams without stepping on them.

• Clear communicator, able to work with engineers and non-technical partners alike.

• Writes things down, and leaves systems easier for others to run than they found them.

• Strong problem-solving skills and attention to detail.

• Commitment to inclusivity, trust, and continuous improvement.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Electronic Arts