C

Cloud Engineer

Chennai, Tamil Nadu, India
5 years exp
and when they do, it’s your responsibility to pick them up and drive them forwar
Posted 19h ago
2 views
Actively Hiring Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Cloud Engineer role at CloudifyOps.

Apply for this position

Apply on Company Website

Job Description

Culture at CloudifyOps : Working at CloudifyOps is a rewarding experience! Great people, a work environment that thrives on creativity, and the opportunity to take on roles beyond a defined

job description

are just some of the reasons you should work with us.

About the Role :

We’re looking for someone who genuinely wants to understand why systems fail, not just respond to alerts. This role sits at the crossroads of cloud infrastructure and production reliability. You’ll own monitoring, handle on-call, and be the person who digs in when things go wrong. At the same time, we’re building an AI-powered pipeline monitoring tool and need someone curious enough to contribute to shaping it, not just watching over it. What you’ll do: Handle the on-call rotation and own incidents end-to-end triage, mitigation, escalation where needed, and clean resolution. You don’t pass the baton and disappear. Write clear, structured RCAs after every significant incident what happened, when, why, and what changes going forward. These go to clients, so they need to work for both an engineer and a non-technical reader. Maintain and improve the monitoring stack across environments dashboards, alerting rules, log pipelines, and distributed traces. Treat noisy alerts as a problem to fix, not something to mute. Provision and manage cloud infrastructure on AWS using Terraform. This is a hands-on role not just reviewing what others set up. Work with Kubernetes across multiple environments, debugging pod and node issues. Monitor CI/CD pipeline health via Jenkins and support teams using Rancher for workload and cluster management. Track application performance using APM tooling and JVM metrics: spot anomalies, investigate degradation, and flag systemic issues before they become incidents. Contribute to the AI monitoring tool initiative: prototype, test, iterate. This is early-stage work and needs someone willing to figure things out, not just execute a finished design. Tech Stack: Cloud & Infrastructure : AWS

  • We’re looking for someone who genuinely wants to understand why systems fail, not just respond to alerts. This role sits at the crossroads of cloud infrastructure and production reliability. You’ll own monitoring, handle on-call, and be the person who digs in when things go wrong. At the same time, we’re building an AI-powered pipeline monitoring tool and need someone curious enough to contribute to shaping it, not just watching over it. What you’ll do: Handle the on-call rotation and own incidents end-to-end triage, mitigation, escalation where needed, and clean resolution. You don’t pass the baton and disappear. Write clear, structured RCAs after every significant incident what happened, when, why, and what changes going forward. These go to clients, so they need to work for both an engineer and a non-technical reader. Maintain and improve the monitoring stack across environments dashboards, alerting rules, log pipelines, and distributed traces. Treat noisy alerts as a problem to fix, not something to mute. Provision and manage cloud infrastructure on AWS using Terraform. This is a hands-on role not just reviewing what others set up. Work with Kubernetes across multiple environments, debugging pod and node issues. Monitor CI/CD pipeline health via Jenkins and support teams using Rancher for workload and cluster management. Track application performance using APM tooling and JVM metrics: spot anomalies, investigate degradation, and flag systemic issues before they become incidents. Contribute to the AI monitoring tool initiative: prototype, test, iterate. This is early-stage work and needs someone willing to figure things out, not just execute a finished design. Tech Stack: Cloud & Infrastructure : AWS
  • Kubernetes (K8s)
  • Terraform
  • Linux Observability & Metrics : Prometheus
  • Grafana
  • APM (Datadog / New Relic / Kfuse)
  • JVM Metrics & GC Analysis
  • ELK / EFK Stack
  • Distributed Tracing CI/CD & Platform : Jenkins
  • ArgoCD
  • Rancher
  • Git
  • Docker Good to Have(Not Mandatory) : Python / Bash scripting
  • OpenTelemetry
  • Zenduty / OpsGenie
  • ML / AI basics Expectations: On-call here is real.Incidents happen outside business hours and when they do, it’s your responsibility to pick them up and drive them forward. That’s not unusual for this type of role but we want to be direct about it upfront. Client expectations are high. You’ll produce RCAs, incident timelines, and status communications that clients read closely. Your writing needs to be clear, structured, and free of vagueness. “We investigated and fixed the issue” isn’t good enough. What was the issue, why did it happen, what was the business impact, and what prevents recurrence. We expect precision regarding your own work. After a change, an incident, or a deployment, you should be a

Frequently Asked Questions

How to apply for Cloud Engineer at CloudifyOps?

Click the "Apply via CareerScan" button on this page.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

5 years of experience is required.

Is this position still open?

Yes, currently active and accepting applications.

Similar Openings

Explore related active roles in Information Technology

View all
Actively Hiring
6–8 years
Salary not disclosed
Saidapet, Tamil Nadu, India
Information TechnologyFull Time
Posted 19h ago
Apply Now
Actively Hiring
13+ years
Salary not disclosed
Saidapet, Tamil Nadu, India
Information TechnologyFull Time
Posted 19h ago
Apply Now
Actively Hiring
0-2 Yrs
Salary not disclosed
Saidapet, Tamil Nadu, India
Information TechnologyFull Time
Posted 19h ago
Apply Now

Cloud Engineer

CloudifyOps · Chennai