Systems Administrator
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
About Business Unit: At the core of all that Epsilon does is a team that sets the foundation of our IT infrastructure.
The team drives innovation and efficiency through pioneering technology across Epsilon's platforms and business verticals.
From being the first point of contact for infrastructure needs to final deployment, the team provides end-to-end solutions for our client-facing platforms.
ETS supports all aspects of revenue-generating platforms for Epsilon and sets the architectural direction for our enterprise deployments.
By adopting the newest technologies, such as Cloud, Automation, and Artificial Intelligence, the team is at the front of redefining our digital business and capturing new opportunities.
Role Summary: The Infrastructure & Reliability Engineer is responsible for operating, supporting, securing, and continuously improving enterprise infrastructure across Windows, Linux, virtualized, cloud, and hybrid environments.
This role combines strong L2 systems administration accountability with foundational L1 Site Reliability Engineering practices, including monitoring, automation, incident response, service health management, and operational improvement.
The role is expected to ensure infrastructure stability, reduce repetitive operational toil, support production reliability, execute changes with control, and collaborate effectively with application, cloud, security, database, network, and vendor teams.
Click here to view how Epsilon transforms marketing with 1 View, 1 Vision and 1 Voice.
Key Skills for This Role
Full Job Posting
Overview
About Business Unit: At the core of all that Epsilon does is a team that sets the foundation of our IT infrastructure.
The team drives innovation and efficiency through pioneering technology across Epsilon's platforms and business verticals.
From being the first point of contact for infrastructure needs to final deployment, the team provides end-to-end solutions for our client-facing platforms.
ETS supports all aspects of revenue-generating platforms for Epsilon and sets the architectural direction for our enterprise deployments.
By adopting the newest technologies, such as Cloud, Automation, and Artificial Intelligence, the team is at the front of redefining our digital business and capturing new opportunities.
Role Summary: The Infrastructure & Reliability Engineer is responsible for operating, supporting, securing, and continuously improving enterprise infrastructure across Windows, Linux, virtualized, cloud, and hybrid environments.
This role combines strong L2 systems administration accountability with foundational L1 Site Reliability Engineering practices, including monitoring, automation, incident response, service health management, and operational improvement.
The role is expected to ensure infrastructure stability, reduce repetitive operational toil, support production reliability, execute changes with control, and collaborate effectively with application, cloud, security, database, network, and vendor teams.
Click here to view how Epsilon transforms marketing with 1 View, 1 Vision and 1 Voice.
Responsibilities
L2 Windows and Linux Infrastructure Operations Administer, troubleshoot, patch, upgrade, and support Windows Server and enterprise Linux platforms across production and non-production environments.
Support Windows services including Active Directory, DNS, Group Policy, WSUS, permissions, access, service accounts, and related infrastructure services.
Perform Linux administration activities covering user management, permissions, file systems, services, packages, processes, logs, networking, storage, and performance troubleshooting.
Support VMware, Proxmox/KVM, physical servers, and cloud-hosted infrastructure across AWS, Azure, or GCP, based on assigned scope.
Execute routine health checks, capacity reviews, backup validation, vulnerability remediation, configuration fixes, and service restoration activities.
Diagnose infrastructure issues across operating systems, storage, compute, network, identity, cloud, virtualization, and security layers.
Support Dell, HPE, and other enterprise server hardware using vendor management interfaces and diagnostic tools where applicable.
Incident, Request, Change, and Problem Management Own and progress incidents, service requests, and change records using ITIL-aligned platforms such as ServiceNow or Jira.
Restore services within agreed SLA targets through structured troubleshooting, escalation, communication, and documentation.
Participate in high-priority incident response, collect evidence, document actions taken, and contribute to RCA and post-incident reviews.
Execute approved changes with proper validation, implementation evidence, rollback planning, and stakeholder communication.
Identify recurring issues and contribute to problem management, permanent fixes, runbook updates, and operational improvements.
Coordinate with application owners, developers, network, cloud, database, security, and vendor teams for investigation and resolution.
L1 Site Reliability Engineering and Observability Monitor availability, performance, service health, capacity, and reliability indicators using approved enterprise monitoring platforms.
Triage alerts, validate impact, follow runbooks, reduce alert noise, and escalate reliability risks to senior SRE, platform, or engineering teams.
Support service reliability reviews by contributing operational metrics, incident trends, alert patterns, and improvement recommendations.
Apply foundational SRE practices such as SLIs, SLOs, error-budget awareness, toil reduction, blameless incident learning, and service health reporting.
Assist with resilience, failover, recovery, and operational-readiness testing under approved plans and senior engineering guidance.
Identify repetitive manual activities and recommend automation, monitoring, documentation, or configuration improvements.
Automation, DevOps, and Platform Improvement Automate routine administration, patching, validation, reporting, deployment, and evidence collection using PowerShell, Bash, Python, Ansible, Terraform, or approved workflow tools.
Use Git-based version control and follow peer review, testing, and release-control practices for scripts, automation, and infrastructure-as-code changes.
Support CI/CD and DevOps operational activities using tools such as Jenkins, GitHub, Bitbucket, Argo CD, Buildkite, or comparable platforms.
Integrate approved REST APIs and workflow tools to improve data consistency, operational efficiency, and task execution quality.
Support basic Docker and Kubernetes operational troubleshooting, including escalation for EKS, GKE, or comparable container platforms where in scope.
Drive continuous improvement by reducing toil, standardizing repeatable tasks, improving runbooks, and increasing operational reliability.
Security, Compliance, and Documentation Implement approved hardening requirements, remediate vulnerabilities, support audit activities, and maintain evidence for operational and security controls.
Maintain accurate SOPs, runbooks, knowledge articles, configuration records, asset updates, change evidence, and troubleshooting documentation.
Follow security policies, access-control requirements, change controls, and compliance processes while supporting infrastructure operations.
Support patch compliance, endpoint protection readiness, backup validation, vulnerability closure, and operational risk reduction.
Success Measures Incidents, service requests, and changes are delivered accurately within agreed service targets.
Windows and Linux platforms remain stable, secure, patched, observable, and documented.
Repetitive operational work is reduced through automation and standardization.
Monitoring alerts are actionable, escalations are timely, and recurring issues are converted into permanent improvement actions.
Changes are implemented safely with validation, rollback planning, evidence, and compliance alignment.
Stakeholders receive clear, timely, and professional communication during operational activities and service-impacting events.
Qualifications
Required Qualifications: 3 to 5 years of hands-on experience in Windows/Linux systems administration, infrastructure operations, production support, cloud operations, DevOps, or SRE.
Strong working knowledge of Windows Server administration and/or Linux administration, with practical exposure to the other platform.
Experience
supporting enterprise infrastructure including virtualization, networking, storage, backup, identity, security, monitoring, and cloud fundamentals.
Practical scripting or automation experience using PowerShell, Bash, Python, Ansible, Terraform, or equivalent tools.
Experience
working with incident, request, change, problem, and knowledge-management processes in an enterprise environment.
Experience
with monitoring, logging, dashboards, alert triage, performance analysis, and capacity management.
Ability to troubleshoot complex infrastructure issues, communicate clearly during service-impacting events, and work with global stakeholders.
Ability to participate in rotational shifts, weekend support, maintenance windows, and on-call support as required.
Preferred Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience. Exposure to AWS, Azure, GCP, VMware, Proxmox/KVM, Docker, Kubernetes, and hybrid infrastructure environments. Foundational understanding of SRE concepts including SLIs, SLOs, error budgets, toil reduction, capacity planning, and blameless post-incident review. Familiarity with observability platforms such as Splunk, ELK/Kibana, Grafana, Prometheus, CloudWatch, Datadog, New Relic, or equivalent. Familiarity with databases such as MySQL or PostgreSQL. Relevant certification in Microsoft, Linux/Red Hat, VMware, AWS, Azure, GCP, ITIL, DevOps, or SRE practices. Core Technical Skills: Operating Systems: Windows Server 2019+, RHEL, CentOS, AlmaLinux, Ubuntu, or comparable Linux distributions. Cloud and Virtualization: AWS, Azure, GCP, VMware vSphere, Proxmox/KVM, physical servers, and hybrid infrastructure. Automation and Infrastructure as Code: PowerShell, Bash, Python, Ansible, Terraform, n8n, or comparable workflow automation tools. DevOps Toolchain: Git, GitHub, Bitbucket, Jenkins, Argo CD, Buildkite, and CI/CD concepts. Observability: Splunk, ELK/Kibana, Grafana, Prometheus, CloudWatch, Datadog, New Relic, or equivalent monitoring platforms. ITSM: ServiceNow, Jira, incident, request, change, problem, and knowledge-management processes. Platform Fundamentals: Active Directory, DNS, Group Policy, patching, storage, backup, APIs, networking, security, capacity, and performance management. Additional Information Epsilon is a global data, technology and services company that powers the marketing and advertising ecosystem. For decades, we’ve provided marketers from the world’s leading brands the data, technology and services they need to engage consumers with 1 View, 1 Vision and 1 Voice. 1 View of their universe of potential buyers. 1 Vision for engaging each individual. And 1 Voice to harmonize engagement across paid, owned and earned channels. Epsilon’s comprehensive portfolio of capabilities across our suite of digital media, messaging and loyalty solutions bridge the divide between marketing and advertising technology. We process 400+ billion consumer actions each day using advanced AI and hold many patents of proprietary technology, including real-time modeling languages and consumer privacy advancements. Thanks to the work of every employee, Epsilon has been consistently recognized as industry-leading by Forrester, Adweek and the MRC. Epsilon is a global company with more than 9,000 employees around the world. Our pillars aren't just words. They're how we show up every day. People centricity: We focus on employee well-being in an environment where colleagues truly care about each other. Collaboration: We work together, support one another, and collectively achieve goals. Growth: There are endless opportunities for growth through learning, development and career advancement. Innovation: We drive progress through cutting-edge solutions and forward-thinking approaches. Flexibility: We’ve created a balance between work and personal life, and we encourage adaptability to solve problems creatively. Our values guide us to create value for our clients, our people and consumers. Act with integrity Work together to win together Innovate with purpose Respect all voices Empower with accountability These pillars and values are our foundation—shaping our culture, guiding our decisions, and uniting us in common purpose. Epsilon is an Equal Opportunity Employer. Epsilon is committed to promoting diversity, inclusion, and equal employment opportunities by using reasonable efforts to attract, recruit, engage and retain qualified individuals of all ethnicities and backgrounds, including, but not limited to, women, people of color, LGBTQ individuals, people with disabilities and any other underrepresented groups, traits or characteristics.
About Epsilon
Parent group profile
Publicis Groupe
A global leader in marketing, communication, and digital transformation.
Visit parent group websiteApply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Epsilon
Product Privacy
Bengaluru, IND
Epsilon is seeking a privacy professional to support privacy operations, program initiatives, and compliance activities. The role covers third-party risk management, data subject requests, data inventory, internal commun
Systems Administrator
Bengaluru, IND
Lead, Client Success
Chicago, USA
Director, Analytic Consulting
Chicago, USA
Senior Lead, Solutions Strategist
Chicago, USA
Integration Project Administrator
Bengaluru, IND
Integration Project Administrator
Bengaluru, IND
Product Privacy
Bengaluru, IND
Client Services Analyst
Westminster, USA