Assume technical leadership for enterprise monitoring and observability initiatives, providing guidance, standards, and best practices across technology teams.
Establish and maintain governance frameworks, operational standards, and monitoring policies to ensure consistent service visibility and alert management throughout the organization.
Develop enterprise monitoring strategies and roadmaps that align with business objectives, operational resiliency goals, infrastructure modernization efforts, and cloud transformation initiatives.
Define monitoring scope, service health objectives, availability metrics, performance standards, and operational requirements for critical business services.
Lead the design and implementation of monitoring solutions supporting infrastructure, cloud platforms, applications, databases, network services, and end-user technologies.
Partner with engineering and operational teams to identify opportunities for increased observability, proactive detection of service degradation, and improved operational efficiency.
Develop and maintain monitoring architecture standards, monitoring templates, dashboard frameworks, alerting policies, and operational runbooks.
Govern alert quality through ongoing optimization of thresholds, event correlation, noise reduction, and root-cause-focused monitoring practices to minimize alert fatigue and improve operational response.
Review operational incidents, service interruptions, monitoring events, and performance trends to identify systemic issues and drive corrective actions that improve service reliability.
Design, implement, and maintain executive and operational dashboards providing visibility into service availability, performance trends, capacity utilization, operational risk, and technology health.
Coordinate enterprise observability initiatives across business and technology teams to ensure monitoring requirements are incorporated into projects, deployments, and service transitions.
Lead the integration of monitoring platforms with enterprise IT Service Management systems, ensuring automated incident creation, escalation workflows, and operational response processes are consistently applied.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
Facilitate governance reviews and operational health assessments to validate compliance with established monitoring standards and service reliability objectives.
Collaborate with application owners, infrastructure teams, and service providers to establish meaningful service-level indicators, operational baselines, and performance benchmarks.
Conduct feasibility studies, technical evaluations, and solution assessments related to monitoring, telemetry collection, event management, observability, automation, and operational analytics.
Evaluate emerging monitoring technologies, Artificial Intelligence for IT Operations (AIOps) capabilities, automation opportunities, and observability platforms to improve service reliability and operational effectiveness.
Establish and maintain strong working relationships with technology leaders, business stakeholders, service providers, engineers, and support organizations.
Develop high-level cost estimates, business cases, and return-on-investment analyses for monitoring enhancements, operational tooling, and technology modernization initiatives.
Present recommendations and cost-benefit analyses related to monitoring platform expansion, service visibility improvements, automation capabilities, and operational maturity efforts.
Ensure monitoring solutions integrate with enterprise architecture, security requirements, operational processes, and service delivery objectives.
Lead enterprise capacity planning initiatives through analysis of monitoring data, utilization trends, performance metrics, and growth projections.
Facilitate monthly operational health and capacity review meetings with infrastructure, cloud, platform engineering, application support, and service management leaders.
Serve as the technical authority for enterprise monitoring platforms like Dynatrace and future observability technologies.
Oversee the creation and maintenance of monitoring templates, notification profiles, event correlation rules, escalation procedures, dashboard standards, and operational reporting frameworks.
Support major incident response activities by providing monitoring expertise, performance analysis, service-impact assessment, and operational visibility during critical situations.
Drive continuous improvement initiatives focused on service reliability, monitoring effectiveness, operational governance, capacity forecasting, and enterprise observability maturity.
Qualifications
6 to 8 years of experience in enterprise systems administration, application administration, infrastructure operations, monitoring platforms, or observability technologies, with an additional 3 years of technical leadership, operational governance, capacity planning, systems design, or project leadership experience preferred.
Demonstrated experience with enterprise monitoring platforms, operational analytics, service health reporting, capacity management, event management, and IT Service Management integrations is highly desirable.
Bachelor's degree in Information Systems, Computer Science, Engineering, or equivalent technical training and experience.
About Caesars Entertainment
Casinos & Gaming50000 employeesFounded 1937
The largest casino and hotel entertainment company in the U.S.