TixelJobs
H
Harnessincvia Greenhouse

Staff Software Engineer – AI SRE

REMOTEPosted 5d ago
OtherStaff+Full-time#remote

Not sure if you're a good fit?

Upload your resume and TixelJobs AI will compare it against Staff Software Engineer – AI SRE at Harnessinc. Get a match score, missing keywords, and improvement tips before you apply.

Free preview · Your resume stays private

About the Role

Harness is the AI Software Delivery Platform company, led by technologist and entrepreneur Jyoti Bansal (founder of AppDynamics, acquired by Cisco for $3.7B). Harness has raised approximately $570M in funding and is valued at $5.5B, backed by leading investors including Goldman Sachs, Menlo Ventures, IVP, Unusual Ventures, Citi Ventures, and more. As AI accelerates code creation, the real bottleneck has shifted to everything after the code – testing, deployments, application security, reliability, compliance, and cost optimization. Harness brings AI and automation to this “outer loop,” helping teams ship software faster while maintaining security and governance throughout the entire software delivery lifecycle.

Powered by Harness AI and the Software Delivery Knowledge Graph, the Harness Platform applies deep context and intelligent automation across the software delivery lifecycle with governance and policy-driven controls embedded throughout the platform. 

Over the past year, Harness powered over 185M deployments, 82M builds, 18T flag evaluations, 8M security scans, 9.1B optimized tests, 3T protected API calls, and helped manage $2.8B in cloud spend — enabling customers like United Airlines, Morningstar, and Choice Hotels to accelerate releases by up to 75%, reduce cloud costs by up to 60%, and achieve 10x DevOps efficiency. 

With a global team across 26 offices and 27 countries, Harness is shaping the future of AI software delivery — and we’re looking for exceptional talent to help us move even faster.

Staff Software Engineer — AI-SRE

About AI SRE

AI is fundamentally changing how engineers build and operate software. At Harness, we're building an AI-native SRE platform that helps engineering teams understand production incidents faster, identify what changed, and automate repetitive operational work.

Our vision goes beyond traditional incident management. We're building AI investigators that can reason across deployments, source code, pull requests, production telemetry, alerts, documentation, runbooks, and organizational knowledge to help engineers answer questions like:

  • What changed?
  • Why did this incident happen?
  • What evidence supports that conclusion?
  • What should we do next?

This role is an opportunity to help define the next generation of AI-powered developer and production operations tooling while solving challenging distributed systems, backend infrastructure, and AI engineering problems.

Why This Role Is Different

Most AI applications today are built around answering questions. We're building systems that investigate, reason, and take action.

Imagine an AI investigator that understands an incident the way a seasoned SRE would — correlating deployments, code changes, feature flags, production telemetry, documentation, previous incidents, and organizational knowledge to determine what likely happened, explain why it happened, and help engineers resolve issues faster.

Building that requires much more than prompting an LLM. It requires designing scalable distributed systems, building intelligent retrieval pipelines, reasoning over complex software delivery data, and creating intuitive developer experiences that engineers trust during high-pressure production incidents.

We're also rethinking how software itself is built. We believe AI will fundamentally change software engineering, and we're looking for engineers who actively experiment with new tools, challenge existing workflows, and help define what an AI-native engineering organization looks like.

About the Role

As a Staff Software Engineer on the AI-SRE team, you will design and build intelligent, scalable platforms that improve service reliability, incident response, and operational efficiency. You will provide technical leadership across the team, own critical components, influence architecture, and collaborate with Site Reliability Engineers and cross-functional teams to solve complex production challenges.

Harness is building an AI-native SRE platform designed to help engineering teams understand production incidents faster by reasoning across deployments, source code, production telemetry, and organizational knowledge. This role is critical for scaling these systems to handle high volumes of alerts and complex production Java problems for large enterprise clients.

What You’ll Do

  • Design, develop, and maintain scalable, highly available AI-SRE platform services.
  • Define technical architecture and author functional specifications and design documents.
  • Own critical system components from design through production operation.
  • Diagnose complex issues across distributed systems and production environments, particularly involving multi-source event processing.
  • Build AI-assisted capabilities for incident detection, diagnosis, remediation, and automation.
  • Establish engineering standards for quality, scalability, security, performance, and reliability.
  • Identify technical debt and scaling risks, then drive improvements across the platform.
  • Design and develop REST, gRPC, GraphQL, and event-driven APIs.
  • Partner with SRE, platform, product, and infrastructure teams during incident investigation.
  • Define observability strategies using metrics, logs, traces, and actionable alerts.
  • Lead technical reviews of architecture, specifications, designs, and code.
  • Mentor Software Engineers and Senior Software Engineers through design reviews and architectural guidance.
  • Influence technical direction across teams while balancing delivery and long-term maintainability.
  • Evaluate emerging AI and platform technologies and apply them to practical reliability problems.

About You

  • Experience: 7 to 10 years of professional software development experience building scalable, distributed applications or platforms.
  • Core Proficiency: Strong experience with Java
  • Leadership: Proven experience leading the architecture and delivery of complex, production-critical systems without relying on direct authority.
  • Technical Depth: Deep understanding of distributed systems, concurrency, resiliency, failure handling, data structures, and algorithms.
  • System Design: Experience designing REST, gRPC, GraphQL, and asynchronous service integrations.
  • Infrastructure: Hands-on experience with Kubernetes, containers, and cloud-native architectures.
  • Operations Mindset: Experience with observability, incident management, and participating in on-call rotations to resolve complex production issues.
  • Execution: A strong bias for execution while maintaining high standards for quality and reliability.

Product Thinking

We value engineers who think beyond implementation.

You should enjoy:

  • Understanding customer problems and thinking from the user's perspective
  • Challenging assumptions and exploring creative solutions
  • Collaborating closely with Product to shape what gets built — not just how it's built
  • Building products that engineers genuinely love using

AI-SRE Specialization

Experience in one or more of the following areas is highly preferred:

  • AIOps & Intelligent Observability: Automated incident response or using telemetry data to detect anomalies and identify root causes.
  • Generative AI: Working with Large Language Models (LLMs), agents, or Retrieval-Augmented Generation (RAG).
  • AI Workflows: Building production AI pipelines with necessary evaluation, monitoring, and safety controls.
  • Automation: Automating operational runbooks and designing human-in-the-loop systems for high-impact production actions.

Technology Stack

Category

Technologies

Languages

Java, Python

Orchestration

Kubernetes, Cloud-native infrastructure

Share