CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Back to all jobs
  1. Home
  2. Jobs
  3. LLM Reliability & Evaluation Engineer
X
XenonStack

LLM Reliability & Evaluation Engineer

Mohali, India
3+ years exp
Full-time
Posted 5d ago
2 views
Actively Hiring Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this LLM Reliability & Evaluation Engineer role at XenonStack.

Apply for this position

Apply on Company Website
Notice a broken link or wrong info?

Job Description

ABOUT XENONSTACK

XenonStack is the fastest-growing

Data and AI Foundry for Agentic Systems

, enabling enterprises to gain

real-time and intelligent business insights

  • We deliver innovation through: -

Agentic Systems for AI Agents

→

Vision AI Platform

→

Inference AI Infrastructure for Agentic Systems

→

Our mission is to accelerate the world’s transition to

AI + Human Intelligence

by making AI agents

reliable, explainable, and enterprise-ready

.


THE OPPORTUNITY

We are seeking an

LLM Reliability & Evaluation Engineer

to ensure that large language models (LLMs) and agentic AI systems meet

enterprise-grade standards of accuracy, safety, and trustworthiness

.

This role focuses on

evaluating, benchmarking, and stress-testing

LLMs in real-world workflows, building frameworks for

reliability, robustness, and continuous improvement

. If you thrive at the intersection of

AI research, applied testing, and responsible deployment

, this is the role for you.


KEY RESPONSIBILITIES

Evaluation Frameworks

  • Design and implement

LLM evaluation pipelines

covering accuracy, robustness, safety, and bias.

  • Develop automated systems for

benchmarking models

on enterprise-relevant tasks.

Reliability Engineering

  • Conduct

stress tests, adversarial testing, and edge-case evaluations

.

  • Build tools to measure

latency, consistency, and error recovery

in multi-turn interactions.

Metrics & Monitoring

  • Define KPIs such as

factual accuracy, hallucination rate, toxicity, and compliance alignment

.

  • Establish real-time monitoring for

drift, anomalies, and performance regressions

.

Collaboration & Alignment

  • Partner with

ML engineers, product managers, and domain experts

to align evaluation with business objectives.

  • Work with Responsible AI teams to implement

ethical, explainable, and compliant evaluation practices

.

Continuous Improvement

  • Feed insights from evaluation into

fine-tuning, RLHF/RLAIF pipelines, and model selection

.

  • Maintain a

central repository of test cases, benchmarks, and evaluation results

.

Research & Innovation

  • Stay current with

state-of-the-art LLM evaluation techniques

, from academic benchmarks to applied enterprise metrics.

  • Explore

automated evaluation using agentic test harnesses and synthetic data generation

.


SKILLS & QUALIFICATIONS

Must-Have

  • 3–6 years in

AI/ML, NLP, or applied model evaluation

.

  • Strong understanding of

LLM architectures, prompt engineering, and failure modes

.

  • Hands-on with

evaluation frameworks

(Eval harnesses, Ragas, OpenAI Evals, DeepEval).

  • Proficiency in

Python

and libraries like

LangChain, LangGraph, LlamaIndex, Hugging Face

.

  • Experience with

vector databases, RAG pipelines, and knowledge graph integration

.

  • Familiarity with

bias/fairness testing and Responsible AI frameworks

.

Good-to-Have

  • Experience with

reinforcement learning (RLHF, RLAIF)

and reward modeling.

  • Exposure to

agentic evaluation frameworks

(multi-agent stress testing, synthetic user simulators).

  • Knowledge of

compliance and safety requirements

for BFSI, GRC, or SOC use cases.

  • Contributions to

open-source evaluation libraries or research papers

.


WHY SHOULD YOU JOIN US?

Agentic AI Product Company

Ensure reliability in cutting-edge AI platforms that are redefining enterprise adoption.
2.

A Fast-Growing Category Leader

Be part of of the fastest-growing

AI Foundries

, powering Fortune 500 enterprises with trustworthy AI.
3.

Career Mobility & Growth

Grow into roles such as

AI Systems Architect, Responsible AI Engineer, or Reliability Engineering Lead

.
4.

Global Exposure

Work on

enterprise-scale evaluation challenges

across BFSI, Healthcare, Telecom, and GRC.
5.

Create Real Impact

Your evaluations will directly shape

production-grade AI agents used in mission-critical systems

.
6.

Culture of Excellence

Our values —

Agency, Taste, Ownership, Mastery, Impatience, and Customer Obsession

— empower you to innovate fearlessly.
7.

Responsible AI First

Join a company that prioritizes

trustworthy, explainable, and compliant AI

.


XENONSTACK CULTURE – JOIN US & MAKE AN IMPACT!

At XenonStack, we believe in

shaping the future of intelligent systems

. We foster a

culture of cultivation

built on bold, human-centric leadership principles, where

deep work, simplicity, and adoption

define everything we do.

Our Cultural Values

Agency

– Be self-directed and proactive.

Taste

– Sweat the details and build with precision.

Ownership

– Take responsibility for outcomes.

Mastery

– Commit to continuous learning and growth.

Impatience

– Move fast and embrace progress.

Customer Obsession

– Always put the customer first.

Our Product Philosophy

Obsessed with Adoption

– Making AI accessible, reliable, and enterprise-ready.

Obsessed with Simplicity

– Turning complex evaluation challenges into seamless, automated frameworks.

Be part of our mission to

accelerate the world’s transition to AI + Human Intelligence

— by making AI agents not just powerful, but

trustworthy and reliable

.

Key Requirements & Skills

  • 3–6 years in AI/ML, NLP, or applied model evaluation.
  • Strong understanding of LLM architectures, prompt engineering, and failure modes.
  • Hands-on with evaluation frameworks (Eval harnesses, Ragas, OpenAI Evals, DeepEval).
  • Proficiency in Python and libraries like LangChain, LangGraph, LlamaIndex, Hugging Face.
  • Experience with vector databases, RAG pipelines, and knowledge graph integration.
  • Familiarity with bias/fairness testing and Responsible AI frameworks.
  • Experience with reinforcement learning (RLHF, RLAIF) and reward modeling.
  • Exposure to agentic evaluation frameworks (multi-agent stress testing, synthetic user simulators).
  • Knowledge of compliance and safety requirements for BFSI, GRC, or SOC use cases.
  • Contributions to open-source evaluation libraries or research papers.

Benefits & Perks

Vision AI Platform** →

Frequently Asked Questions

How to apply for LLM Reliability & Evaluation Engineer at XenonStack?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

3+ years of experience is required.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Broken link or expired?
X

XenonStack

Visit Company Website

More jobs at XenonStack

Talent Acquisition Specialist

Mohali, India

Senior Software Engineer, Quality

Mohali, India

Data Engineer II

Mohali, India

Share this Opening

Job Alerts for ai_engineering

Receive email alerts whenever new ai_engineering roles in Mohali are posted.

Set Free Alert →

Similar Openings

Explore related active roles in ai_engineering

View all
UrgentActively Hiring
Finc
Senior AI Engineer
Finc Verified
0-2 Yrs
Salary not disclosed
Bengaluru
ai_engineeringFull-time
Posted 1d ago
Apply Now
UrgentActively Hiring
Fa-espx-saasfaprod1
Senior AI Engineer
Fa-espx-saasfaprod1 Verified
8+ years
Salary not disclosed
Pune, Maharashtra, India
ai_engineeringFull-time
Posted 1d ago
Apply Now
UrgentActively Hiring
Grafana Labs
Staff AI Engineer - Grafana AI/ML | Canada | Remote
Grafana Labs Verified
0-2 Yrs
Salary not disclosed
Canada (Remote)
ai_engineeringFull-timeRemote
Posted 1d ago
Apply Now

LLM Reliability & Evaluation Engineer

XenonStack · Mohali

Apply on Company Website