CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Back to all jobs
  1. Home
  2. Jobs
  3. Lead Machine Learning Engineer - ML Infrastructure
jobgether
jobgether

Lead Machine Learning Engineer - ML Infrastructure

Canada
₹13.6L/mo
10+ years exp
Full-time
Posted 2d ago
0 views
Actively Hiring Urgent Opening Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Lead Machine Learning Engineer - ML Infrastructure role at jobgether.

Apply for this position

Apply on Company Website
Notice a broken link or wrong info?

Job Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Machine Learning Engineer - ML Infrastructure based in the Canada .

This role serves as the technical anchor for machine learning infrastructure supporting large-scale AI products and applications.
You will own the architecture and evolution of an end-to-end ML platform spanning training, experimentation, inference, and edge deployment.
The position connects applied machine learning, data platform, security, firmware, and infrastructure teams around shared technical direction.
You will make high-impact architectural decisions that influence multiple product teams and production systems at significant scale.
The role combines hands-on engineering with technical leadership, reliability, security, developer experience, and operational ownership.
You will help deliver ML-powered capabilities from initial design through production deployment while balancing performance, latency, cost, and research velocity.
This is a senior technical leadership opportunity where your decisions can directly influence real-world safety, efficiency, and customer outcomes.

Accountabilities:

  • Set the technical strategy and own end-to-end delivery of the machine learning platform across training, experimentation, batch and inference, and edge deployment.

  • Make architectural decisions across ML infrastructure layers and serve as the primary technical accountability point for multiple AI product teams.

  • Design, launch, and continuously improve ML-powered features while co-owning production outcomes such as safety metrics, reliability, performance, and cost.

  • Design and operate scalable and batch inference systems using technologies such as

Ray and Spark

, including deployment patterns, observability, service-level objectives (SLOs), and unified training-to-production workflows.

  • Partner with firmware and edge engineering teams to package, validate, and deploy machine learning models to connected devices.

  • Build feedback loops between edge deployments and cloud infrastructure to support continuous model and system improvement.

  • Own reliability, observability, security, and operational practices for ML systems spanning cloud and edge environments.

  • Establish and improve on-call practices, incident response processes, infrastructure hardening, and production reliability standards.

  • Own or co-own end-to-end technical delivery for high-priority and high-risk initiatives, from modeling and system architecture through production rollout.

  • Serve as the technical authority for ML infrastructure architecture and establish direction for applied ML, firmware, security, and data platform teams.

  • Mentor senior engineers and applied scientists while helping teams make sound technical trade-offs at the appropriate level of abstraction.

  • Improve developer experience through documentation, engineering standards, reusable practices, and clear platform guidance.

  • Contribute to and represent the organization within relevant open-source communities, including

Ray, Spark, RayDP, and Kubernetes

.

  • Balance research velocity with platform stability and communicate technical trade-offs effectively across science and engineering teams.

  • Champion customer-focused, long-term, inclusive, collaborative, and growth-oriented engineering practices.

Requirements

10+ years of experience in machine learning engineering

, with demonstrated technical leadership across at least two major ML platform domains such as distributed training, data or research infrastructure, cloud inference, or feature engineering.

  • Proven track record of delivering ML-powered products or features end-to-end, from technical design through production deployment and iteration, with measurable product or business impact.

  • Strong hands-on expertise with

Ray and Kubernetes

in production environments.

  • Strong experience with

Spark

is highly preferred.

  • Deep understanding of machine learning fundamentals beyond pipelines, including evaluation methodology, dataset design, ablation, model drift, and the ability to review and redirect modeling approaches.

  • Ability to bridge research and engineering teams and translate ML requirements into scalable, reliable production systems.

  • Demonstrated cross-organizational technical leadership, including influencing platform decisions, roadmaps, and go/no-go decisions based on throughput, latency, reliability, and cost trade-offs.

  • Strong understanding of distributed ML infrastructure and production systems at scale.

  • Experience navigating the trade-offs between scientific experimentation and platform stability, with the communication skills needed to align both sides.

  • Strong architectural judgment and ability to operate as the senior technical authority for complex ML infrastructure initiatives.

  • Excellent communication and collaboration skills, with the ability to influence senior technical stakeholders without relying solely on formal authority.

  • Strong mentoring and technical leadership capabilities.

  • Prior contributions to open-source projects such as

Ray, Spark, RayDP, or Kubernetes

are a plus.

  • Experience with enterprise security and compliance requirements in ML environments is a plus.

  • Experience with edge or on-device machine learning and collaboration with firmware or embedded engineering teams is a plus.

Benefits

  • Annual base salary range of

$196,000–$269,500 CAD

, with actual compensation varying based on factors such as location, knowledge, skills, and experience.

  • Eligibility for an initial

RSU grant with no vesting cliff

, subject to applicable plan terms.

  • Ongoing equity refresh opportunities tied to performance, subject to plan terms and conditions.

  • Performance-based bonus or variable compensation for eligible roles.

  • Flexible, employee-led remote working model.

  • Comprehensive health benefits.

  • Parental leave programs.

  • Professional development stipend.

  • Opportunities to work on high-impact AI and machine learning infrastructure at significant scale.

  • Opportunity to influence platform architecture and technical direction across multiple product teams.

  • Exposure to cloud, edge, distributed computing, machine learning, and open-source technologies.

  • Collaborative environment emphasizing long-term ownership, growth, inclusion, and customer outcomes.

  • Reasonable accommodations available throughout the recruitment process for qualified candidates.

  • How Jobgether works:
    We use an

AI-powered matching process

to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
https://jobgether.com/how-jobgether-works  Why Apply Through Jobgether?   

  • Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

 
 
#LI-CL1

Key Requirements & Skills

10+ years of experience in machine learning engineering

, with demonstrated technical leadership across at least two major ML platform domains such as distributed training, data or research infrastructure, cloud inference, or feature engineering.

  • Proven track record of delivering ML-powered products or features end-to-end, from technical design through production deployment and iteration, with measurable product or business impact.
  • Strong hands-on expertise with

Ray and Kubernetes

in production environments.

  • Strong experience with

Spark

is highly preferred.

  • Deep understanding of machine learning fundamentals beyond pipelines, including evaluation methodology, dataset design, ablation, model drift, and the ability to review and redirect modeling approaches.
  • Ability to bridge research and engineering teams and translate ML requirements into scalable, reliable production systems.
  • Demonstrated cross-organizational technical leadership, including influencing platform decisions, roadmaps, and go/no-go decisions based on throughput, latency, reliability, and cost trade-offs.
  • Strong understanding of distributed ML infrastructure and production systems at scale.
  • Experience navigating the trade-offs between scientific experimentation and platform stability, with the communication skills needed to align both sides.
  • Strong architectural judgment and ability to operate as the senior technical authority for complex ML infrastructure initiatives.
  • Excellent communication and collaboration skills, with the ability to influence senior technical stakeholders without relying solely on formal authority.
  • Strong mentoring and technical leadership capabilities.
  • Prior contributions to open-source projects such as

Ray, Spark, RayDP, or Kubernetes

are a plus.

  • Experience with enterprise security and compliance requirements in ML environments is a plus.
  • Experience with edge or on-device machine learning and collaboration with firmware or embedded engineering teams is a plus.

Benefits & Perks

Benefits

  • Annual base salary range of

$196,000–$269,500 CAD

, with actual compensation varying based on factors such as location, knowledge, skills, and experience.

Frequently Asked Questions

How to apply for Lead Machine Learning Engineer - ML Infrastructure at jobgether?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

The salary for this role is $196,000 per annum.

What experience is required?

10+ years of experience is required.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Broken link or expired?
jobgether

jobgether

Looking for a work-from-home job? Get matched with the best flexible and remote jobs in the world with Jobgether.

Visit Company Website

More jobs at jobgether

Sr. Sales Specialist, Govt. HighQ & Partnership Sales

US

Sr. Symitar Systems Analyst

US

Staff Backend Engineer

Canada

Share this Opening

Job Alerts for devops

Receive email alerts whenever new devops roles in Canada are posted.

Set Free Alert →

Similar Openings

Explore related active roles in devops

View all
Actively Hiring
Finc
Platform Engineer
Finc Verified
7+ years
Salary not disclosed
Atlanta
devopsFull-time
Posted 18h ago
Apply Now
Actively Hiring
Supabase
AI Platform Engineer
Supabase Verified
0-2 Yrs
₹7/mo
Remote, Global
devopsFull-timeRemote
Posted 18h ago
Apply Now
UrgentActively Hiring
Eurofins
IT Infrastructure Service Desk Agent
Eurofins Verified
3+ years
Salary not disclosed
Coimbatore, TN, in
devopsFull-time
Posted 18h ago
Apply Now

Lead Machine Learning Engineer - ML Infrastructure

jobgether · Canada

Apply on Company Website