Scan your resume against ATS criteria for this Site Reliability Engineer (SRE) role at Amplifire.
Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. Working alongside DevOps engineers within the Platform Operations team, this role combines software engineering and systems operations with a focus on observability, automation, incident reduction, and operational excellence.
You will partner closely with software engineering, QA, security, and DevOps to establish reliability practices, improve production visibility, strengthen incident response, automate operational workflows, and help engineering teams deliver changes safely and confidently.
This role supports systems operating under regulatory and compliance requirements, including FedRAMP and SOC 2. The ideal candidate understands that reliability, traceability, security, and change management enable sustainable development velocity rather than compete with it.
Amplifire expects all technical team members to leverage AI-assisted tools and workflows as force multipliers for productivity, learning, automation, and problem-solving while maintaining strong engineering skills, sound judgment, security standards, and operational accountability.
· Establish and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and other measures of system health.
· Build and continuously improve monitoring, logging, tracing, dashboards, and alerting that provide actionable visibility into application and infrastructure health.
· Analyze system behavior, performance, capacity, and reliability trends to proactively identify risks and improvement opportunities.
· Partner with engineering teams to define reliability requirements and improve the availability, scalability, and performance of production systems.
· Participate in the on-call rotation and respond to production incidents with urgency and sound technical judgment.
· Improve incident detection, triage, escalation, communication, mitigation, and recovery processes.
· Lead or contribute to blameless post-incident reviews, root cause analysis, and corrective actions.
· Develop and maintain runbooks, recovery procedures, and operational practices that improve service resilience and reduce mean time to detect and restore service.
· Identify and eliminate operational toil through automation, self-service tooling, and continuous improvement.
· Build and maintain cloud infrastructure using Infrastructure as Code practices and tools such as Terraform or AWS CDK.
· Improve CI/CD pipelines, deployment safeguards, rollback capabilities, and progressive delivery practices.
· Develop internal tools and automation that improve reliability, resilience, and engineering productivity.
· Support scalable, secure, and cost-effective production environments.
· Partner with software engineers to embed reliability, operability, and observability throughout the development lifecycle while reducing operational friction and helping teams safely own their services in production.
· Help engineering teams diagnose complex production issues across application and infrastructure layers.
· Contribute to security, compliance, capacity-planning, cloud cost optimization, and operational standards.
· Use AI-assisted tools responsibly to improve troubleshooting, automation, documentation, and operational efficiency while sharing knowledge and continuously improving team practices.
How to apply for Site Reliability Engineer (SRE) at Amplifire?
Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.
What is the salary for this role?
Salary details will be discussed during the interview.
What experience is required?
This position is open to freshers and experienced candidates.
Is this position still open?
Yes, currently active and accepting applications.
Site Reliability Engineer (SRE)
Amplifire · India