CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Back to all jobs
  1. Home
  2. Jobs
  3. Senior Site Reliability Engineer
MyOperator
MyOperator

Senior Site Reliability Engineer

Noida, India
3+ years exp
Full-time
Posted 5d ago
2 views
Actively Hiring Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Senior Site Reliability Engineer role at MyOperator.

Apply for this position

Apply on Company Website
Notice a broken link or wrong info?

Job Description

About MyOperator

MyOperator is a Business AI Operator platform that enables businesses, teams, and AI agents to work together seamlessly for customer operations such as Sales, Support, Escalations, Feedback, and Refund processes. With 12,000+ businesses using our platform, we operate at meaningful scale and power mission-critical communication workflows including voice bots, WhatsApp automation, and intelligent call routing.

We are building for reliability, speed, and impact. MyOperator values ownership, critical thinking, and execution. This is a high-expectation, high-learning environment where engineers are empowered to solve complex problems and build systems that directly affect customer outcomes.

Role Overview

We are looking for a skilled and proactive Site Reliability Engineer (SRE) to take end-to-end ownership of production reliability, observability, and performance engineering across MyOperator’s AI-powered communication infrastructure.

This role is not operational-only — it requires strong system design thinking, deep troubleshooting ability, and a production ownership mindset. You will define reliability standards, build observability frameworks, lead incident response, and drive SLO-based engineering practices across distributed AWS and Kubernetes environments.

Key Responsibilities

  • Own production reliability, uptime, latency, and error budgets across critical services.
  • Design and manage production-grade monitoring using Grafana, VictoriaMetrics (Prometheus), and AWS CloudWatch.
  • Define and enforce SLIs, SLOs, and SLA thresholds for AI communication systems (voice bots, WhatsApp APIs, call routing).
  • Build real-time operational dashboards for incident response, capacity planning, and leadership visibility.
  • Implement end-to-end distributed tracing using OpenTelemetry (OTEL Collector).
  • Design and maintain centralized logging with strong correlation between logs, metrics, and traces.
  • Create SLO-based alerting systems with minimal noise and fast incident detection.
  • Lead incident response lifecycle: alert triage, mitigation, RCA documentation, and preventive improvements.
  • Drive MTTR reduction through structured monitoring, automation, and reliability engineering practices.
  • Monitor and troubleshoot AWS EKS (Kubernetes) production workloads.
  • Instrument and monitor LLM API integrations, AI inference pipelines, and messaging systems.
  • Analyze logs using OpenSearch / ELK for anomaly detection and root cause identification.
  • Automate operational workflows using Python or Bash to eliminate manual toil.
  • Drive performance optimization, scalability improvements, and capacity planning.
  • Collaborate with engineering teams to instrument new services from day.

Required Skills & Qualifications

  • 3–6 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
  • Hands-on experience with:

VictoriaMetrics / Prometheus (time-series monitoring)

Grafana dashboards and visualization

PromQL for writing complex queries and alerts

  • Experience implementing distributed tracing using OpenTelemetry (Mandatory).
  • Strong experience with centralized logging systems (ELK / OpenSearch / Loki).
  • Experience with alerting frameworks such as Alertmanager or Grafana Alerts.
  • Strong understanding of SLIs, SLOs, SLA design, and reliability engineering principles.
  • Hands-on experience managing AWS production workloads (EC2, RDS, ELB, CloudWatch, IAM).
  • Experience with Kubernetes (AWS EKS preferred).
  • Good understanding of Linux systems, networking, and cloud infrastructure.
  • Experience handling production incidents and participating in on-call rotations.
  • Ability to automate operational tasks using Python or Bash.

Good to Have

  • Experience with OpenSearch / ELK log pipelines and anomaly detection.
  • Kubernetes monitoring (pod health, node metrics, autoscaling behavior).
  • CI/CD observability integration (Jenkins, GitHub Actions).
  • Experience monitoring LLM APIs and AI inference pipelines.
  • Familiarity with MLOps or AI observability tools (Arize, WhyLabs, etc.).
  • Service mesh exposure (Istio).
  • Infrastructure as Code (Terraform, CloudFormation).
  • Experience with chaos engineering or load testing tools.
  • Multi-cluster or multi-region architecture exposure.

Key Expectations

  • Ownership of production systems and high availability.
  • Strong troubleshooting and debugging skills.
  • Focus on automation and reliability improvements.
  • Proactive approach to incident prevention.
  • Ability to reduce alert noise and improve signal quality.
  • Data-driven approach to reliability engineering.

This Role Is Not For

  • Candidates with purely development experience and no production ownership.
  • Candidates without real incident response or on-call experience.
  • Freshers or candidates with less than 3 years of experience.

Key Requirements & Skills

  • 3–6 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
  • Hands-on experience with:
  • Experience with OpenSearch / ELK log pipelines and anomaly detection.
  • Kubernetes monitoring (pod health, node metrics, autoscaling behavior).
  • CI/CD observability integration (Jenkins, GitHub Actions).
  • Experience monitoring LLM APIs and AI inference pipelines.
  • Familiarity with MLOps or AI observability tools (Arize, WhyLabs, etc.).
  • Service mesh exposure (Istio).
  • Infrastructure as Code (Terraform, CloudFormation).
  • Experience with chaos engineering or load testing tools.
  • Multi-cluster or multi-region architecture exposure.

Frequently Asked Questions

How to apply for Senior Site Reliability Engineer at MyOperator?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

3+ years of experience is required.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Broken link or expired?
MyOperator

MyOperator

About MyOperator – The Business AI Operator Talk To Sales: +91 92129 92129 [email protected] Login Pricing Platform Business AI Operator AI Suite WhatsApp Marketing Suite Call Management Suite Business AI Operator Automate business operations with a unified AI system AI Suite Intelligent AI chatbots and voicebots for omnichannel communication WhatsApp Marketing Suite Official WhatsApp Marketing Suite for unified customer communication Call Management Suite Smart IVR, routing, and call management with WhatsApp sync Unified Analytics & Reporting Track performances and metrics for a multi-

Visit Company Website

More jobs at MyOperator

Front Deployed Engineer

Noida, India

Associate Vice President - Strategic Alliances & Partnerships

Noida, India

Intern Partner Acquisition

Noida, India

Share this Opening

Job Alerts for sre

Receive email alerts whenever new sre roles in Noida are posted.

Set Free Alert →

Similar Openings

Explore related active roles in sre

View all
UrgentActively Hiring
Icertis
Software Engineer, Cloud Site Reliability (SRE)
Icertis Verified
0-2 Yrs
Salary not disclosed
Pune, Maharashtra, India
sreFull-time
Posted 5d ago
Apply Now
UrgentActively Hiring
BNY
Senior Vice President, Site Reliability Engineer Manager
BNY Verified
0-2 Yrs
Salary not disclosed
Pune, MH, India
sreFull-time
Posted 5d ago
Apply Now
Actively Hiring
JP Morgan Chase
Lead Site Reliability Engineer + AWS
JP Morgan Chase Verified
5+ years
₹2.9L – ₹5.8L/mo
Bengaluru, Karnataka, India
sreFull-time
Posted 5d ago
Apply Now

Senior Site Reliability Engineer

MyOperator · Noida

Apply on Company Website