Staff Software Engineer - AI SRE
Software Engineering, Data Science
Bengaluru, Karnataka, India
Harness is the AI Software Delivery Platform company, led by technologist and entrepreneur Jyoti Bansal (founder of AppDynamics, acquired by Cisco for $3.7B). Harness has raised approximately $570M in funding and is valued at $5.5B, backed by leading investors including Goldman Sachs, Menlo Ventures, IVP, Unusual Ventures, Citi Ventures, and more. As AI accelerates code creation, the real bottleneck has shifted to everything after the code – testing, deployments, application security, reliability, compliance, and cost optimization. Harness brings AI and automation to this “outer loop,” helping teams ship software faster while maintaining security and governance throughout the entire software delivery lifecycle.
Powered by Harness AI and the Software Delivery Knowledge Graph, the Harness Platform applies deep context and intelligent automation across the software delivery lifecycle with governance and policy-driven controls embedded throughout the platform.
Over the past year, Harness powered over 185M deployments, 82M builds, 18T flag evaluations, 8M security scans, 9.1B optimized tests, 3T protected API calls, and helped manage $2.8B in cloud spend — enabling customers like United Airlines, Morningstar, and Choice Hotels to accelerate releases by up to 75%, reduce cloud costs by up to 60%, and achieve 10x DevOps efficiency.
With a global team across 26 offices and 27 countries, Harness is shaping the future of AI software delivery — and we’re looking for exceptional talent to help us move even faster.
Staff Software Engineer — AI-SRE
About AI SRE
AI is fundamentally changing how engineers build and operate software. At Harness, we're building an AI-native SRE platform that helps engineering teams understand production incidents faster, identify what changed, and automate repetitive operational work.
Our vision goes beyond traditional incident management. We're building AI investigators that can reason across deployments, source code, pull requests, production telemetry, alerts, documentation, runbooks, and organizational knowledge to help engineers answer questions like:
- What changed?
- Why did this incident happen?
- What evidence supports that conclusion?
- What should we do next?
This role is an opportunity to help define the next generation of AI-powered developer and production operations tooling while solving challenging distributed systems, backend infrastructure, and AI engineering problems.
Why This Role Is Different
Most AI applications today are built around answering questions. We're building systems that investigate, reason, and take action.
Imagine an AI investigator that understands an incident the way a seasoned SRE would — correlating deployments, code changes, feature flags, production telemetry, documentation, previous incidents, and organizational knowledge to determine what likely happened, explain why it happened, and help engineers resolve issues faster.
Building that requires much more than prompting an LLM. It requires designing scalable distributed systems, building intelligent retrieval pipelines, reasoning over complex software delivery data, and creating intuitive developer experiences that engineers trust during high-pressure production incidents.
We're also rethinking how software itself is built. We believe AI will fundamentally change software engineering, and we're looking for engineers who actively experiment with new tools, challenge existing workflows, and help define what an AI-native engineering organization looks like.
About the Role
As a Staff Software Engineer on the AI-SRE team, you will design and build intelligent, scalable platforms that improve service reliability, incident response, and operational efficiency. You will provide technical leadership across the team, own critical components, influence architecture, and collaborate with Site Reliability Engineers and cross-functional teams to solve complex production challenges.
Harness is building an AI-native SRE platform designed to help engineering teams understand production incidents faster by reasoning across deployments, source code, production telemetry, and organizational knowledge. This role is critical for scaling these systems to handle high volumes of alerts and complex production Java problems for large enterprise clients.
What You’ll Do
- Design, develop, and maintain scalable, highly available AI-SRE platform services.
- Define technical architecture and author functional specifications and design documents.
- Own critical system components from design through production operation.
- Diagnose complex issues across distributed systems and production environments, particularly involving multi-source event processing.
- Build AI-assisted capabilities for incident detection, diagnosis, remediation, and automation.
- Establish engineering standards for quality, scalability, security, performance, and reliability.
- Identify technical debt and scaling risks, then drive improvements across the platform.
- Design and develop REST, gRPC, GraphQL, and event-driven APIs.
- Partner with SRE, platform, product, and infrastructure teams during incident investigation.
- Define observability strategies using metrics, logs, traces, and actionable alerts.
- Lead technical reviews of architecture, specifications, designs, and code.
- Mentor Software Engineers and Senior Software Engineers through design reviews and architectural guidance.
- Influence technical direction across teams while balancing delivery and long-term maintainability.
- Evaluate emerging AI and platform technologies and apply them to practical reliability problems.
About You
- Experience: 7 to 10 years of professional software development experience building scalable, distributed applications or platforms.
- Core Proficiency: Strong experience with Java
- Leadership: Proven experience leading the architecture and delivery of complex, production-critical systems without relying on direct authority.
- Technical Depth: Deep understanding of distributed systems, concurrency, resiliency, failure handling, data structures, and algorithms.
- System Design: Experience designing REST, gRPC, GraphQL, and asynchronous service integrations.
- Infrastructure: Hands-on experience with Kubernetes, containers, and cloud-native architectures.
- Operations Mindset: Experience with observability, incident management, and participating in on-call rotations to resolve complex production issues.
- Execution: A strong bias for execution while maintaining high standards for quality and reliability.
Product Thinking
We value engineers who think beyond implementation.
You should enjoy:
- Understanding customer problems and thinking from the user's perspective
- Challenging assumptions and exploring creative solutions
- Collaborating closely with Product to shape what gets built — not just how it's built
- Building products that engineers genuinely love using
AI-SRE Specialization
Experience in one or more of the following areas is highly preferred:
- AIOps & Intelligent Observability: Automated incident response or using telemetry data to detect anomalies and identify root causes.
- Generative AI: Working with Large Language Models (LLMs), agents, or Retrieval-Augmented Generation (RAG).
- AI Workflows: Building production AI pipelines with necessary evaluation, monitoring, and safety controls.
- Automation: Automating operational runbooks and designing human-in-the-loop systems for high-impact production actions.
Technology Stack
Category | Technologies |
|---|---|
Languages | Java, Python |
Orchestration | Kubernetes, Cloud-native infrastructure |
Databases | MongoDB, PostgreSQL, TimescaleDB, Vector databases |
APIs & Integration | REST, gRPC, GraphQL, Event-driven systems |
Cloud | Google Cloud Platform (GCP), AWS, or Azure |
Observability | Metrics, logs, traces, and Cloud Monitoring |
Technical Competencies
- In-depth tactical knowledge of incident management, on-call orchestration, AI-driven root cause analysis, and SLO/SLI frameworks — paired with broad experience across distributed systems, event-driven architectures, and real-time collaboration platforms
- Domain and product expert when representing Harness AI-SRE to customers, prospects, and internal teams alike — fluent in the full incident lifecycle from alert ingestion through post-mortem automation
- Seamlessly include, promote, and balance cross-functional inputs from Product, Design, and Enablement to shape features spanning the alert-to-resolution journey
- Break down complex technical requirements — such as multi-source event processing, AI investigator pipelines, and integration frameworks — to motivate and lead cross-functional teams
- Consider the execution needed today while making architectural investments aligned with broader impact — from shared service directories and pipeline steps to proactive AI capabilities and enterprise-grade RBAC
Preferred Qualifications
- Experience with AWS, Azure, or Google Cloud Platform.
- Experience in incident management, on-call orchestration, SLO/SLI frameworks, incident management practices and AI-assisted root cause analysis.
- Experience building internal developer platforms or reliability tooling.
- Customer-facing experience representing technical products to enterprise stakeholders.
- Strong architectural judgment balancing near-term delivery with long-term scalability, extensibility, and enterprise security.
- Experience designing the complete incident lifecycle—from alert ingestion through resolution and post-mortem automation.
- Ability to translate complex requirements into scalable AI investigator, event-processing, and integration platforms.
- Bachelor’s degree in Computer Science or a related discipline; an advanced degree is preferred.
- Equivalent professional experience will also be considered.
What Success Looks Like
- Delivering reliable AI-SRE capabilities that measurably reduce operational effort.
- Improving incident detection, diagnosis, and recovery times.
- Increasing platform scalability, observability, and maintainability.
- Raising engineering quality through architecture, standards, reviews, and mentorship.
- Enabling teams to operate production services more safely and efficiently.
Harness in the news:
- Accelerating Our Mission to Bring AI to Everything After Code
- Goldman Sachs leads investment in software delivery startup Harness at $5.5 billion valuation
- How Harness runs 16 “startups within a startup” at scale | Jyoti Bansal
- Harness Research Shows AI Visibility Crisis Fueling Security Nightmare
- Harness has been named to the Inc. Power Partner list for software delivery success
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex or national origin.
At Harness, we care about your privacy and are committed to protecting your personal data. For additional information on this topic, you can visit our privacy Portal: https://harness-privacy.relyance.ai/
Note on Fraudulent Recruiting/Offers
We have become aware that there may be fraudulent recruiting attempts being made by people posing as representatives of Harness. These scams may involve fake job postings, unsolicited emails, or messages claiming to be from our recruiters or hiring managers.
Please note, we do not ask for sensitive or financial information via chat, text, or social media, and any email communications will come from the domain @harness.io. Additionally, Harness will never ask for any payment, fee to be paid, or purchases to be made by a job applicant. All applicants are encouraged to apply directly to our open jobs via our website. Interviews are generally conducted via Zoom video conference unless the candidate requests other accommodations.
If you believe that you have been the target of an interview/offer scam by someone posing as a representative of Harness, please do not provide any personal or financial information and contact us immediately at security@harness.io. You can also find additional information about this type of scam and report any fraudulent employment offers via the Federal Trade Commission’s website (https://consumer.ftc.gov/articles/job-scams), or you can contact your local law enforcement agency.