Scan your resume against ATS criteria for this Lead Data & AI Ops Engineer, Data & AI Engineering role at pretiumenterpriseservices.
We are seeking a high-potential, hands-on
to own and continuously improve the operational health, governance, controls, reliability, and efficiency of our enterprise Data & AI ecosystem.
This is a high-impact technical leadership role with
. The successful candidate will establish the operating model, engineering controls, automation, observability, and governance required to run Data & AI platforms as reliable, secure, and cost-efficient enterprise services.
The ideal candidate combines deep
with a strong operations and controls mindset. This engineer will also
required to achieve operational excellence and efficiency goals.
· Supported
across enterprise data platforms, data pipelines, analytics, BI, and AI/ML workloads, ensuring availability, reliability, performance, and SLA adherence.
· Provided day-to-day
, including workload monitoring, query performance analysis, troubleshooting, access/RBAC management, capacity monitoring, and platform health checks.
· Supported and enhanced
, troubleshooting data ingestion, transformation, orchestration, processing, and downstream data delivery issues across production environments.
· Supported
, including monitoring application and model-related jobs, data dependencies, API integrations, scheduled processes, failures, and overall operational health.
· Contributed to
capabilities by using AI/GenAI tools for incident analysis, log summarization, anomaly identification, troubleshooting assistance, knowledge retrieval, and faster root-cause analysis.
· Developed
to reduce repetitive operational activities, automate health checks and validations, accelerate issue resolution, and improve support productivity.
· Supported the implementation of
across data pipelines, Snowflake workloads, and AI services to proactively identify failures, performance degradation, unusual patterns, and operational risks.
· Assisted in developing
for common production issues, reducing manual intervention and improving the Resolution SLA.
· Used
to support troubleshooting, incident summarization, RCA preparation, log analysis, runbook recommendations, and knowledge management activities.
· Monitored production
, investigated failures, performed impact analysis, and coordinated timely service restoration.
· Performed
to identify data discrepancies and ensure accurate, complete, and reliable data delivery to downstream applications and AI/analytics workloads.
· Supported enterprise data platform controls covering
.
· Monitored
, identified inefficient queries and workloads, and supported optimization initiatives to improve performance and control platform costs.
· Built and maintained
across Data and AI platforms.
· Managed
activities, including production troubleshooting, service restoration, RCA documentation, change validation, deployment support, and permanent remediation of recurring issues.
· Supported
, including CI/CD pipelines, testing, deployment, release validation, version control, monitoring, documentation, and production support.
· Worked closely with
to troubleshoot production issues, manage dependencies, and implement platform improvements.
· Participated in
, ensuring critical Data and AI incidents were addressed within agreed SLAs and appropriately communicated to stakeholders.
· Identified recurring operational issues and implemented
to reduce manual effort, prevent repeat incidents, and improve production stability.
· Contributed to continuous improvement by promoting
.
·
across Data Engineering, Data Platforms, Data Ops, Cloud Engineering, or Production Operations, with demonstrated technical leadership.
·
, including architecture, administration, SQL, performance tuning, workload management, security/RBAC, monitoring, troubleshooting, and optimization.
· Strong experience
, not just administering or supporting Data Platforms.
· Demonstrated
, with measurable outcomes in Snowflake/cloud consumption reduction, workload optimization, cost attribution, and efficiency improvement.
· Strong experience building automation using
.
· Strong expertise with
and enterprise ETL/ELT technologies such as Fivetran, Informatica, and Azure Data Factory.
· Experience implementing
.
· Strong understanding of production operations, incident/problem management, RCA, change management, and platform reliability engineering.
· Experience with enterprise BI platforms such as
.
· Ability to operate as both a
, taking problems from identification through solution architecture, engineering, implementation, and measurable business outcome.
· Prefer candidates already residing in Bangalore.
· Standard Shift Timing is 12noon to 9pm, however this may vary depending on the business requirements.
· 3 Days work from office.
· Weekend on call support is required.
How to apply for Lead Data & AI Ops Engineer, Data & AI Engineering at pretiumenterpriseservices?
Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.
What is the salary for this role?
Salary details will be discussed during the interview.
What experience is required?
12+ years of experience is required.
Is this position still open?
Yes, currently active and accepting applications.
Explore related active roles in ai_engineering
Lead Data & AI Ops Engineer, Data & AI Engineering
pretiumenterpriseservices · Bangalore