ABOUT XENONSTACK
XenonStack is the fastest-growing
Data and AI Foundry for Agentic Systems
, enabling enterprises to gain
real-time and intelligent business insights
- We deliver innovation through: -
Agentic Systems for AI Agents
→
Vision AI Platform
→
Inference AI Infrastructure for Agentic Systems
→
Our mission is to accelerate the world’s transition to
AI + Human Intelligence
by making AI agents
reliable, explainable, and enterprise-ready
.
THE OPPORTUNITY
We are seeking an
LLM Reliability & Evaluation Engineer
to ensure that large language models (LLMs) and agentic AI systems meet
enterprise-grade standards of accuracy, safety, and trustworthiness
.
This role focuses on
evaluating, benchmarking, and stress-testing
LLMs in real-world workflows, building frameworks for
reliability, robustness, and continuous improvement
. If you thrive at the intersection of
AI research, applied testing, and responsible deployment
, this is the role for you.
KEY RESPONSIBILITIES
Evaluation Frameworks
LLM evaluation pipelines
covering accuracy, robustness, safety, and bias.
- Develop automated systems for
benchmarking models
on enterprise-relevant tasks.
Reliability Engineering
stress tests, adversarial testing, and edge-case evaluations
.
latency, consistency, and error recovery
in multi-turn interactions.
Metrics & Monitoring
factual accuracy, hallucination rate, toxicity, and compliance alignment
.
- Establish real-time monitoring for
drift, anomalies, and performance regressions
.
Collaboration & Alignment
ML engineers, product managers, and domain experts
to align evaluation with business objectives.
- Work with Responsible AI teams to implement
ethical, explainable, and compliant evaluation practices
.
Continuous Improvement
- Feed insights from evaluation into
fine-tuning, RLHF/RLAIF pipelines, and model selection
.
central repository of test cases, benchmarks, and evaluation results
.
Research & Innovation
state-of-the-art LLM evaluation techniques
, from academic benchmarks to applied enterprise metrics.
automated evaluation using agentic test harnesses and synthetic data generation
.
SKILLS & QUALIFICATIONS
Must-Have
AI/ML, NLP, or applied model evaluation
.
LLM architectures, prompt engineering, and failure modes
.
evaluation frameworks
(Eval harnesses, Ragas, OpenAI Evals, DeepEval).
Python
and libraries like
LangChain, LangGraph, LlamaIndex, Hugging Face
.
vector databases, RAG pipelines, and knowledge graph integration
.
bias/fairness testing and Responsible AI frameworks
.
Good-to-Have
reinforcement learning (RLHF, RLAIF)
and reward modeling.
agentic evaluation frameworks
(multi-agent stress testing, synthetic user simulators).
compliance and safety requirements
for BFSI, GRC, or SOC use cases.
open-source evaluation libraries or research papers
.
WHY SHOULD YOU JOIN US?
Agentic AI Product Company
Ensure reliability in cutting-edge AI platforms that are redefining enterprise adoption.
2.
A Fast-Growing Category Leader
Be part of of the fastest-growing
AI Foundries
, powering Fortune 500 enterprises with trustworthy AI.
3.
Career Mobility & Growth
Grow into roles such as
AI Systems Architect, Responsible AI Engineer, or Reliability Engineering Lead
.
4.
Global Exposure
Work on
enterprise-scale evaluation challenges
across BFSI, Healthcare, Telecom, and GRC.
5.
Create Real Impact
Your evaluations will directly shape
production-grade AI agents used in mission-critical systems
.
6.
Culture of Excellence
Our values —
Agency, Taste, Ownership, Mastery, Impatience, and Customer Obsession
— empower you to innovate fearlessly.
7.
Responsible AI First
Join a company that prioritizes
trustworthy, explainable, and compliant AI
.
XENONSTACK CULTURE – JOIN US & MAKE AN IMPACT!
At XenonStack, we believe in
shaping the future of intelligent systems
. We foster a
culture of cultivation
built on bold, human-centric leadership principles, where
deep work, simplicity, and adoption
define everything we do.
Our Cultural Values
Agency
– Be self-directed and proactive.
Taste
– Sweat the details and build with precision.
Ownership
– Take responsibility for outcomes.
Mastery
– Commit to continuous learning and growth.
Impatience
– Move fast and embrace progress.
Customer Obsession
– Always put the customer first.
Our Product Philosophy
Obsessed with Adoption
– Making AI accessible, reliable, and enterprise-ready.
Obsessed with Simplicity
– Turning complex evaluation challenges into seamless, automated frameworks.
Be part of our mission to
accelerate the world’s transition to AI + Human Intelligence
— by making AI agents not just powerful, but
trustworthy and reliable
.