Base Career helps you apply smarter for this job.
Key skills for this role
At EA, we believe games are powerful because they bring together multiple ways people engage: play, watch, create, and connect.
And increasingly, the biggest entertainment platforms aren't just places to consume content — they're places where communities build.
Creator-made content is already a proven part of EA's history — from community creation tools in Battlefield to The Gallery in The Sims 4.
We believe new creative technologies and tools will expand how players engage with and contribute to our experiences, supported by thoughtful product design, safety systems, and global reach.
Our focus is on enabling more players to participate in creative expression by making creation easier, safer, and more rewarding.
As Lead Infrastructure Engineer, you will own the GPU fleet our researchers train on including capacity, scheduling, diagnostics, and support.
You will set technical direction for GPU operations and infrastructure architecture.
You will additionally lead Infrastructure as Code setup, granting permissions, and debugging infrastructure problems.
This is a hybrid role, working three days per week in Redwood City, Montreal, or Vancouver.
You will report to the Head of Data and Infrastructure.
You will own GPU fleet operations across our AWS estate.
You will build the scheduling layer from zero.
You will diagnose GPU and node failures fast and completely and drive hardware evidence and replacement through AWS support and capacity-block channels.
You will run researcher support as a first-class product including holding office hours, owning the support channel, and driving the recurring causes out of existence with self-service tooling, preflight checks, and documentation
You will instrument the fleet including utilization, queue depth, job success rate, and cost per experiment metrics.
You will partner with our external compute and lab partnerships as a technical contact, and with EA's central infrastructure groups on shared services and escalation.
You will author runbooks, decision records, and onboarding docs.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Montréal, CAN
Kirkland, USA
Montréal, CAN
Vancouver, CAN
Vancouver, CAN
Vancouver, CAN
Redwood City, USA
Vancouver, CAN
Vancouver, CAN
8+ years of experience operating production infrastructure, with deep, current, hands-on AWS depth — EC2 GPU fleets, EKS, IAM and cross-account security, VPC and networking, S3 and FSx
Experience scheduling, diagnosing, and managing GPUs in AWS specifically
Experience operating GPU fleets at 1000+ GPU scale
Expertise in scripting and automation with Python, PowerShell, bash, or equivalent
Expertise in infrastructure as code (Terraform or equivalent)
Familiarity with a GPU scheduling or orchestration layer (like Slurm, Kubernetes with Kueue or Volcano, Ray, dStack or SkyPilot)
Observability practice including Grafana, Prometheus, or equivalent
Electronic Arts is a global video game company that develops, publishes, and operates interactive entertainment franchises across sports, action, racing, simulation, and online services.
Visit company websiteJobs and hiring trendsLead · 8+ years experience
Hybrid
Apply faster on company sites with our extension.