{bc}
greenhouse

Staff Software Engineer – AI SRE

Harness
Bengaluru, IND
Senior · 10+ years experience
Discovered 1 weeks ago
awsazuregcpgraphqlgrpcjava
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

awsazuregcp
Smart Apply

Full Job Posting

About AI SRE

AI is fundamentally changing how engineers build and operate software. At Harness, we're building an AI-native SRE platform that helps engineering teams understand production incidents faster, identify what changed, and automate repetitive operational work.

Our vision goes beyond traditional incident management. We're building AI investigators that can reason across deployments, source code, pull requests, production telemetry, alerts, documentation, runbooks, and organizational knowledge to help engineers answer questions like:

What changed?

Why did this incident happen?

What evidence supports that conclusion?

What should we do next?

This role is an opportunity to help define the next generation of AI-powered developer and production operations tooling while solving challenging distributed systems, backend infrastructure, and AI engineering problems.

Why This Role Is Different

Most AI applications today are built around answering questions. We're building systems that investigate, reason, and take action.

Imagine an AI investigator that understands an incident the way a seasoned SRE would — correlating deployments, code changes, feature flags, production telemetry, documentation, previous incidents, and organizational knowledge to determine what likely happened, explain why it happened, and help engineers resolve issues faster.

Building that requires much more than prompting an LLM. It requires designing scalable distributed systems, building intelligent retrieval pipelines, reasoning over complex software delivery data, and creating intuitive developer experiences that engineers trust during high-pressure production incidents.

We're also rethinking how software itself is built. We believe AI will fundamentally change software engineering, and we're looking for engineers who actively experiment with new tools, challenge existing workflows, and help define what an AI-native engineering organization looks like.

About the Role

As a Staff Software Engineer on the AI-SRE team, you will design and build intelligent, scalable platforms that improve service reliability, incident response, and operational efficiency. You will provide technical leadership across the team, own critical components, influence architecture, and collaborate with Site Reliability Engineers and cross-functional teams to solve complex production challenges.

Harness is building an AI-native SRE platform designed to help engineering teams understand production incidents faster by reasoning across deployments, source code, production telemetry, and organizational knowledge. This role is critical for scaling these systems to handle high volumes of alerts and complex production Java problems for large enterprise clients.

What You’ll Do

Design, develop, and maintain scalable, highly available AI-SRE platform services.

Define technical architecture and author functional specifications and design documents.

Own critical system components from design through production operation.

Diagnose complex issues across distributed systems and production environments, particularly involving multi-source event processing.

Build AI-assisted capabilities for incident detection, diagnosis, remediation, and automation.

Establish engineering standards for quality, scalability, security, performance, and reliability.

Identify technical debt and scaling risks, then drive improvements across the platform.

Design and develop REST, gRPC, GraphQL, and event-driven APIs.

Partner with SRE, platform, product, and infrastructure teams during incident investigation.

Define observability strategies using metrics, logs, traces, and actionable alerts.

Lead technical reviews of architecture, specifications, designs, and code.

Mentor Software Engineers and Senior Software Engineers through design reviews and architectural guidance.

Influence technical direction across teams while balancing delivery and long-term maintainability.

Evaluate emerging AI and platform technologies and apply them to practical reliability problems.

About You

Experience: 7 to 10 years of professional software development experience building scalable, distributed applications or platforms.

Core Proficiency: Strong experience with Java

Leadership: Proven experience leading the architecture and delivery of complex, production-critical systems without relying on direct authority.

Technical Depth: Deep understanding of distributed systems, concurrency, resiliency, failure handling, data structures, and algorithms.

System Design: Experience designing REST, gRPC, GraphQL, and asynchronous service integrations.

Infrastructure: Hands-on experience with Kubernetes, containers, and cloud-native architectures.

Operations Mindset: Experience with observability, incident management, and participating in on-call rotations to resolve complex production issues.

Execution: A strong bias for execution while maintaining high standards for quality and reliability.

Product Thinking

We value engineers who think beyond implementation.

You should enjoy:

Understanding customer problems and thinking from the user's perspective

Challenging assumptions and exploring creative solutions

Collaborating closely with Product to shape what gets built — not just how it's built

Building products that engineers genuinely love using

AI-SRE Specialization

Experience in one or more of the following areas is highly preferred:

AIOps & Intelligent Observability: Automated incident response or using telemetry data to detect anomalies and identify root causes.

Generative AI: Working with Large Language Models (LLMs), agents, or Retrieval-Augmented Generation (RAG).

AI Workflows: Building production AI pipelines with necessary evaluation, monitoring, and safety controls.

Automation: Automating operational runbooks and designing human-in-the-loop systems for high-impact production actions.

Technology Stack

Category

Technologies

Languages

Java, Python

Orchestration

Kubernetes, Cloud-native infrastructure

Databases

MongoDB, PostgreSQL, TimescaleDB, Vector databases

APIs & Integration

REST, gRPC, GraphQL, Event-driven systems

Cloud

Google Cloud Platform (GCP), AWS, or Azure

Observability

Metrics, logs, traces, and Cloud Monitoring

Technical Competencies

In-depth tactical knowledge of incident management, on-call orchestration, AI-driven root cause analysis, and SLO/SLI frameworks — paired with broad experience across distributed systems, event-driven architectures, and real-time collaboration platforms

Domain and product expert when representing Harness AI-SRE to customers, prospects, and internal teams alike — fluent in the full incident lifecycle from alert ingestion through post-mortem automation

Seamlessly include, promote, and balance cross-functional inputs from Product, Design, and Enablement to shape features spanning the alert-to-resolution journey

Break down complex technical requirements — such as multi-source event processing, AI investigator pipelines, and integration frameworks — to motivate and lead cross-functional teams

Consider the execution needed today while making architectural investments aligned with broader impact — from shared service directories and pipeline steps to proactive AI capabilities and enterprise-grade RBAC

Preferred Qualifications

Experience with AWS, Azure, or Google Cloud Platform.

Experience in incident management, on-call orchestration, SLO/SLI frameworks, incident management practices and AI-assisted root cause analysis.

Experience building internal developer platforms or reliability tooling.

Customer-facing experience representing technical products to enterprise stakeholders.

Strong architectural judgment balancing near-term delivery with long-term scalability, extensibility, and enterprise security.

Experience designing the complete incident lifecycle—from alert ingestion through resolution and post-mortem automation.

Ability to translate complex requirements into scalable AI investigator, event-processing, and integration platforms.

Bachelor’s degree in Computer Science or a related discipline; an advanced degree is preferred.

Equivalent professional experience will also be considered.

What Success Looks Like

Delivering reliable AI-SRE capabilities that measurably reduce operational effort.

Improving incident detection, diagnosis, and recovery times.

Increasing platform scalability, observability, and maintainability.

Raising engineering quality through architecture, standards, reviews, and mentorship.

Enabling teams to operate production services more safely and efficiently.

Harness in the news:

Accelerating Our Mission to Bring AI to Everything After Code

Goldman Sachs leads investment in software delivery startup Harness at $5.5 billion valuation

How Harness runs 16 “startups within a startup” at scale | Jyoti Bansal

Harness Research Shows AI Visibility Crisis Fueling Security Nightmare

Harness has been named to the Inc. Power Partner list for software delivery success

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex or national origin.

At Harness, we care about your privacy and are committed to protecting your personal data. For additional information on this topic, you can visit our privacy Portal: https://harness-privacy.relyance.ai/

Note on Fraudulent Recruiting/Offers

We have become aware that there may be fraudulent recruiting attempts being made by people posing as representatives of Harness. These scams may involve fake job postings, unsolicited emails, or messages claiming to be from our recruiters or hiring managers.

Please note, we do not ask for sensitive or financial information via chat, text, or social media, and any email communications will come from the domain @harness.io. Additionally, Harness will never ask for any payment, fee to be paid, or purchases to be made by a job applicant. All applicants are encouraged to apply directly to our open jobs via our website. Interviews are generally conducted via Zoom video conference unless the candidate requests other accommodations.

If you believe that you have been the target of an interview/offer scam by someone posing as a representative of Harness, please do not provide any personal or financial information and contact us immediately at security@harness.io. You can also find additional information about this type of scam and report any fraudulent employment offers via the Federal Trade Commission’s website ( https://consumer.ftc.gov/articles/job-scams), or you can contact your local law enforcement agency.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Harness