{bc}
company_site

Développeur.se - opérations liées à l’apprentissage automatique / MLOps Engineer

Electronic Arts
Montréal, CAN
Senior · 8+ years experience
Hybrid
Discovered Today
awseksgrafanaprometheuspythonpytorch
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

awseksgrafana
Smart Apply

Full Job Posting

« Pour visualiser la description de poste en français, veuillez sélectionner le français dans le menu déroulant au haut de la page. »

At EA, we believe games are powerful because they bring together multiple ways people engage: play, watch, create, and connect.

And increasingly, the biggest entertainment platforms aren't just places to consume content — they're places where communities build.

Creator-made content is already a proven part of EA's history — from community creation tools in Battlefield to The Gallery in The Sims 4.

We believe new creative technologies and tools will expand how players engage with and contribute to our experiences, supported by thoughtful product design, safety systems, and global reach.

Our focus is on enabling more players to participate in creative expression by making creation easier, safer, and more rewarding.

As an MLOps Engineer, you will work closely with researchers to optimize models, help shape the data they need for training, and instrument the system to measure and improve model performance.

Working in close partnership with the Lead Infrastructure Engineer and Research Leads, you will translate requirements into durable operational processes and ensure the fleet operates effectively.

You will report to the Head of Data & Infrastructure.

Responsibilities

You will optimize how our models train and serve.

You will build the streaming data path from lake to GPU.

You will make our data formats a deliberate choice.

You will instrument all of it including Per-job GPU utilization, queue depth and job success rate, cost per experiment, data-loader stall time, and serving latency percentiles.

You will land the training and experiment telemetry into one place researchers and planners trust.

Qualifications

8+ years of engineering experience

4+ years building and operating ML systems in production

Expertise in Python

Comfort reading and modifying framework-level code — PyTorch, and an inference server such as vLLM, SGLang or TensorRT-LLM

Demonstrated model or inference optimization work

Hands-on experience with distributed data processing for ML (including Ray, Spark, or equivalent) and columnar and ML-native formats (such as Parquet, Arrow, Lance, or similar)

Practical observability skills including Prometheus, Grafana, or equivalent

Working knowledge of AWS compute and storage as they apply to ML workloads — EC2 GPU instances, S3 and its performance tiers, EKS

Experience with containerization and CI/CD for ML artifacts and a reproducibility instinct

This is a hybrid role, onsite three days a week, in Redwood City, Montreal, or Vancouver.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Electronic Arts