CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Back to all jobs
  1. Home
  2. Jobs
  3. Site Reliability Engineer (SRE)
T
ttecdigital

Site Reliability Engineer (SRE)

Hyderabad
8+ yrs exp
Posted 21 Sept 2026
2 views
Actively Hiring

Check Your Resume Match Score

Scan your resume against ATS criteria for this Site Reliability Engineer (SRE) role at ttecdigital.

Apply for this position

Apply on Company Website
Notice a broken link or wrong info?

Job Description

At TTEC Digital, we coach clients to ensure their employees feel valued, and fully supported, because an amazing customer experience is an employee first process. Our vision is the same, a place where employees know they can thrive

  • The role: Own production reliability for a real-time platform where uptime and latency ARE the product — voice, desktop, intelligence, and AI combined; an agent mid-call can't wait for a retry.
    First SRE hired immediately (Day 0–14) for production scaling and SLO ownership; a second joins at the start of Phase 3 for 24/7 coverage.
    Pairs with C1 Platform Foundation on observability and tenancy isolation
  • Startup environment: weekly deploys, 1-week sprints, fail fast, move forward — reliability engineering at that speed, not against it. 
    What you'll own:
    SLOs and error budgets per tenant/service
    Incident response and blameless postmortems
    Production scaling and capacity
    Observability depth (p50/p95/p99 per event hop)
    Uptime as a personal mission
    On-call rotation with DevOps
    Your committed timelines
  • Who you are: Self-starter, grit, show-me mentality — you prove reliability with dashboards and drills, not assertions.
    A ways-to-YES engineer: weekly deploys are the heartbeat and your job is making them safe, never slowing them.
    You love new technology, adapt fast when the stack changes under you, use AI tools daily to multiply velocity, and consider yourself exceptional.
    Calm in an incident, relentless after it.
    Team player who likes winning.
    8+ years operating production systems at scale; owns SLOs, error budgets, incident command.
    Strong Go or Python — you automate reliability, you don't toil at it
  • Everything you build is code: runbooks execute, remediation is automatic, toil trends to zero.
    Deep on event-driven and real-time systems reliability — NATS-class buses, WebSocket fleets, streaming pipelines — and the failure physics underneath: state, race conditions, locking, ordering, back-pressure, cascading load. You've debugged these in production.
    Strong monitoring and uptime mindset — metrics, logs, traces wired to alerting that catches it before the customer does; you know the difference between a noisy alert and a real signal.
    Good networking understanding — protocols and how they work (TCP/UDP, TLS, WebSocket, DNS, load balancing); RTP/SIP a strong plus for our media paths.
    GCP at scale; multi-cloud literacy a plus. Multi-tenancy isolation experience a strong plus.
    Capacity modeling and load testing partnership with QA — find the knee of the curve before customers do.
    Chaos engineering — failure injection as routine practice; prove graceful degradation, don't assume it.
    Deploy-safety partnership with DevOps — canary analysis, automatic rollback triggers, error-budget-driven release gates.
    AI-aware reliability — monitoring model latency, drift, and cost as production signals, not just CPU and memory.
    Incident communication craft — clear, fast, blameless; execs and customers get truth at the right altitude.
    A master debugger of production — reads the trace, the metric, the flame graph, and sees it; narrows an incident to the service, the deploy, the event
  • What You Will Bring: 8+ years operating production systems at scale; owns SLOs, error budgets, incident command. 
    Strong Go or Python — you automate reliability, you don't toil at it
  • Everything you build is code: runbooks execute, remediation is automatic, toil trends to zero. 
    Deep on event-driven and real-time systems reliability — NATS-class buses, WebSocket fleets, streaming pipelines — and the failure physics underneath: state, race conditions, locking, ordering, back-pressure, cascading load. You've debugged these in production. 
    Strong monitoring and uptime mindset — metrics, logs, traces wired to alerting that catches it before the customer does; you know the difference between a noisy alert and a real signal. 
    Good networking understanding — protocols and how they work (TCP/UDP, TLS, WebSocket, DNS, load balancing); RTP/SIP a strong plus for our media paths. 
    GCP at scale; multi-cloud literacy a plus. Multi-tenancy isolation experience a strong plus. 
    Capacity modeling and load testing partnership with QA — find the knee of the curve before customers do. 
    Chaos engineering — failure injection as routine practice; prove graceful degradation, don't assume it. 
    Deploy-safety partnership with DevOps — canary analysis, automatic rollback triggers, error-budget-driven release gates. 
    AI-aware reliability — monitoring model latency, drift, and cost as production signals, not just CPU and memory. 
    Incident communication craft — clear, fast, blameless; execs and customers get truth at the right altitude. 
    A master debugger of production — reads the trace, the metric, the flame graph, and sees it; narrows an incident to the service, the deploy, the event.

Frequently Asked Questions

How to apply for Site Reliability Engineer (SRE) at ttecdigital?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

8+ yrs of experience is required.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Broken link or expired?
T

ttecdigital

More jobs at ttecdigital

Agentic AI Architect - – Google Gemini Enterprise CX

Austin, TX

Amazon Connect and Agentic AI Specialist, Senior Principal Consultant

Austin, TX

Amazon Connect CX and Platform Specialist, Senior Consultant

Austin, TX

Share this Opening

Job Alerts for Software Engineering

Receive email alerts whenever new Software Engineering roles in Hyderabad are posted.

Set Free Alert →

Similar Openings

Explore related active roles in Software Engineering

View all
Actively Hiring
YouTrip
IT System Administrator
YouTrip Verified
0-2 Yrs
Salary not disclosed
Chennai
Software EngineeringFull Time
Posted 21 Sept 2026
Apply Now
Actively Hiring
zetaglobal
Lead Backend Engineer
zetaglobal Verified
8-12 yrs
Salary not disclosed
Bangalore, IND
Software EngineeringFull Time
Posted 21 Sept 2026
Apply Now
Actively Hiring
zetaglobal
Lead Full Stack Engineer
zetaglobal Verified
8+ yrs
Salary not disclosed
Bangalore, IND
Software EngineeringFull Time
Posted 21 Sept 2026
Apply Now

Site Reliability Engineer (SRE)

ttecdigital · Hyderabad

Apply on Company Website