Base Career helps you apply smarter for this job.
Key skills for this role
We are seeking a customer-focused Support Engineer to drive incident resolution, root-cause analysis (RCA), and performance optimization across Workday’s enterprise platform and autonomous AI agent workflows. In this high-visibility role, you will analyze system metrics, debug cloud-hosted ML service pipelines, inspect LLM orchestration layers, and manage critical customer escalations within strict SLAs. You will also partner directly with engineering and data science teams through feature iteration and optimization. A key part of this role involves hands-on AI evaluation: analyzing LLM outputs, reviewing conversation logs, and digging into system traces to spot failure modes and translate those insights into prompt, data, and workflow improvements.
Key Responsibilities
Enterprise SaaS & Functional Domain Expertise: Apply operational knowledge of enterprise applications and workflows to validate AI logic and troubleshoot functional processing errors.
Hands-On AI Evaluation: Regularly review LLM outputs, AI conversation logs, and execution traces to identify edge cases, hallucinations, and failure modes. Perform data labeling and translate diagnostic insights into actionable updates for prompts, workflows, and system logic.
Technical Troubleshooting & RCA: Perform root-cause analysis on software defects, performance bottlenecks, and LLM agent execution failures using Kibana, Grafana, and other cloud telemetry tools.
Cloud & LLM Diagnostics: Debug enterprise AI workflows hosted across public cloud environments (AWS, GCP), isolating issues across model hosting services, API gateways, and LLM reasoning pipelines.
Incident & Queue Management: Triage high-severity (P1) support queues, enforce SLAs, and prioritize critical outages over routine inquiries. Participate in weekend on-call rotations for continuous coverage.
Customer Success: Act as the primary technical escalation point for customer IT leadership, clearly explaining root causes, workarounds, and resolution plans during critical incidents.
Database & Code Diagnostics: Write complex SQL queries to validate backend data integrity, debug REST/SOAP API payloads (JSON/XML), and use Python or Bash scripts to automate diagnostics.
Reliability & Product Partnership: Document detailed investigation traces in Jira, ServiceNow, or Salesforce, update runbooks, and partner with engineering and data science teams to deliver permanent fixes and address issue trends.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Vancouver, CAN
Your work days are brighter here. We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and a
Pune, IND
Pune, IND
Pleasanton, USA
Your work days are brighter here. We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and a
Pleasanton, USA
Your work days are brighter here. We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and a
, USA
Your work days are brighter here. We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and a
, USA
Your work days are brighter here. We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and a
Vancouver, CAN
Pune, IND
Pune, IND
Pune, IND
Pleasanton, USA
Pleasanton, USA
, USA
, USA
We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you’ll feel it. Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We’re in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you’ll do meaningful work with Workmates who’ve got your back. In return, we’ll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you’ve found a match in Workday, and we hope to be a match for you too.
We are seeking a customer-focused Support Engineer to drive incident resolution, root-cause analysis (RCA), and performance optimization across Workday’s enterprise platform and autonomous AI agent workflows. In this high-visibility role, you will analyze system metrics, debug cloud-hosted ML service pipelines, inspect LLM orchestration layers, and manage critical customer escalations within strict SLAs. You will also partner directly with engineering and data science teams through feature iteration and optimization. A key part of this role involves hands-on AI evaluation: analyzing LLM outputs, reviewing conversation logs, and digging into system traces to spot failure modes and translate those insights into prompt, data, and workflow improvements.
Key Responsibilities
Enterprise SaaS & Functional Domain Expertise: Apply operational knowledge of enterprise applications and workflows to validate AI logic and troubleshoot functional processing errors.
Hands-On AI Evaluation: Regularly review LLM outputs, AI conversation logs, and execution traces to identify edge cases, hallucinations, and failure modes. Perform data labeling and translate diagnostic insights into actionable updates for prompts, workflows, and system logic.
Technical Troubleshooting & RCA: Perform root-cause analysis on software defects, performance bottlenecks, and LLM agent execution failures using Kibana, Grafana, and other cloud telemetry tools.
Cloud & LLM Diagnostics: Debug enterprise AI workflows hosted across public cloud environments (AWS, GCP), isolating issues across model hosting services, API gateways, and LLM reasoning pipelines.
Incident & Queue Management: Triage high-severity (P1) support queues, enforce SLAs, and prioritize critical outages over routine inquiries. Participate in weekend on-call rotations for continuous coverage.
Customer Success: Act as the primary technical escalation point for customer IT leadership, clearly explaining root causes, workarounds, and resolution plans during critical incidents.
Database & Code Diagnostics: Write complex SQL queries to validate backend data integrity, debug REST/SOAP API payloads (JSON/XML), and use Python or Bash scripts to automate diagnostics.
Reliability & Product Partnership: Document detailed investigation traces in Jira, ServiceNow, or Salesforce, update runbooks, and partner with engineering and data science teams to deliver permanent fixes and address issue trends.
Enterprise cloud applications for finance and human resources.
Visit company websiteFull-time
Mid · 3+ years experience
Hybrid
Apply faster on company sites with our extension.