CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Export All Job Postings
Back to all jobs
  1. Home
  2. Jobs
  3. Applied AI Research Engineer: Benchmarking & Performance Economics
A
Akka

Applied AI Research Engineer: Benchmarking & Performance Economics

San Francisco, California
Full-time
Posted 5d ago
0 views
Skills:AIAPIAccounting/FinanceAgentic AIAutoGen+15 more
Actively Hiring Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Applied AI Research Engineer: Benchmarking & Performance Economics role at Akka.

Apply for this position

Apply on Company Website
Notice a broken link or wrong info?

Job Description

This is an applied research role, not a machine learning science role. You will not be advancing the state of the art in model architecture. You will be producing decision-grade numbers: benchmarks, cost models, and calculators that determine what we build, what we tell customers, and what we are willing to claim in public.

The work sits at the intersection of three disciplines that rarely overlap in person — systems engineering, inference-stack depth, and honest experimental design. If you have ever read a vendor benchmark and immediately knew which confound made it meaningless, this is the seat for you.-

The token economics of agentic execution.

What it actually costs to run an autonomous agent to task completion, and how that compares to LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK, Ray, and whatever exists six months from now. Tokens per call is the wrong unit and you know it, the unit is cost per completed task at a fixed success rate, and getting there means measuring steps, retries, tail latency, success rate, and run-to-run variance.

Routing and model-selection economics.

Where each routing strategy sits on the cost/quality frontier, what the router costs to run, and how a 5% misroute rate compounds into a 40% task failure rate across a thirty-step trajectory. Routing quality gets evaluated at the trajectory level here, not the request level.

Small language model economics

  • The crossover analysis: at what volume, task narrowness, and quality tolerance does a fine-tuned 3B beat a frontier API call
  • You own the full stack: data curation, training, evaluation, serving, and the cost of drift and retraining, and the judgment to say "don't train this" when the numbers say so.

GPU and serving-stack modeling.

A parameterized calculator that takes model size, quantization, batch size, context distribution, KV cache footprint, and concurrency, and returns throughput, memory headroom, latency percentiles, and cost per million tokens. Validated against measured ground truth, not derived from spec sheets. It accounts for prefix caching, because agentic workloads resend the same system prompt hundreds of times and a model that ignores that is wrong by a multiple.

The cost of governance.

What policy evaluation, guardrails, evaluator calls, and human-in-the-loop suspension actually cost in latency and tokens. Enterprise buyers assume the number is bad. Being able to state it precisely turns an objection into a differentiator.

The benchmark harness as a product.

Pinned versions, controlled cache state, variance reported rather than averaged away, full configuration captured with every result, reproducible by a third party. Most benchmark work in this industry does not survive scrutiny. Yours will be the asset that does.

Mindset (the part we care most about)

You are constitutionally unwilling to publish a number you can't defend.

When the result is inconvenient, you report the result.

Fair to the competition, to the point of discomfort.

You will implement a rival framework properly, in its own idiom, and you will spend the extra two days doing it well. An unfair benchmark is worse than no benchmark, it's a liability the moment someone reproduces it.

You have been wrong in public and corrected it.

We consider this a qualification, not a blemish. Someone who has never had a result challenged has never had a result that mattered.

Scoping discipline.

Most of this work arrives as an ill-posed question. Turning "is our routing better?" into a measurable experiment with a defined success criterion is the core skill, and it happens before any code is written.

Suspicious of your own instruments.

You assume the harness is lying until you've proven otherwise. You measure the measurement.

Writer.

The output of this role is arguments supported by evidence. A correct result that can't be explained to an engineer, an architect, and a buyer has delivered a fraction of its value.

Builder, not administrator.

Akka is small enough that you write this playbook rather than inherit it. That should energize you.

Experience (the part we'll flex on for the right mindset)

Real experimental design and statistics.

Confidence intervals, confound control, sample sizing, and the judgment to know when a difference is noise. This is the most common gap in otherwise strong candidates and the we're least able to flex on.

Inference stack depth, hands-on.

vLLM, SGLang, or TensorRT-LLM in anger. Quantization formats and their quality trade-offs. Continuous batching, paged attention, prefix caching, GPU memory arithmetic.

Strong systems engineering.

Python for the harness, plus enough comfort in a systems language to read and profile serving code. What you build here is infrastructure others extend. It needs to be maintainable, not notebook-shaped.

Evaluation design for non-deterministic systems.

LLM-as-judge and its failure modes, task-completion rubrics, trajectory-level evaluation. You know that a benchmark measuring the wrong thing precisely is worse than measuring the right thing roughly.

Cost modeling that survives interrogation.

Comfort building models a finance-literate stakeholder will go through line by line: amortization, utilization assumptions, marginal versus fully-loaded cost.

Nice to have, not required

  • Hands-on time across multiple agent frameworks, sufficient to build the same task idiomatically in each.
  • Distributed systems intuition, concurrency, backpressure, failure modes, tail latency behavior.
  • Published benchmarks, teardowns, or analyses under your own name that held up to challenge.
  • Familiarity with event-driven architectures or actor-model platforms.

Explicitly not required

A PhD. A publication record. Novel architecture or training-methods research. Deep theoretical ML. We are not screening for any of these, and candidates should not self-select out for lacking them.

How we'll evaluate

We weigh rigor and approach above years-of-experience checkboxes. The interview loop is built to surface:

A benchmark you have actually built

(published or internal) and how you handled variance, what confounds you controlled for, and what you got wrong.

  • A work sample: we hand you a specific performance or cost claim from a competitor's marketing page and ask you to design the experiment that would confirm or refute it in a day. We are watching for how fast you find the unstated assumptions and how honestly you scope what you'd leave unmeasured.

How you reason about fairness

when the comparison makes us look worse than we'd like.

How you'd explain your most technical result

to someone evaluating our platform against three alternatives.
A candidate with a strong research pedigree and no hands-on serving experience is not a fit for this seat, regardless of institution. A candidate who has spent three years profiling inference workloads, has shipped a fine-tuned model, and can explain exactly why their last benchmark was subtly wrong probably is.

  • Competitive salary with performance-based incentives.
  • Comprehensive health and wellness benefits.
  • Opportunities for professional development and continuous learning.
  • Flexible remote working environment.
  • Collaborative, inclusive, and innovative company culture.
  • A transparent, distributed work environment with a strong focus on

work-life balance

.

  • Challenging work that interacts with innovative applications used by millions.
  • A collaborative culture that attracts the "brightest minds" in the technology community.

Key Skills & Requirements

20 identified

Click any skill to discover matching job openings across India:

AIAPIAccounting/FinanceAgentic AIAutoGenCrewAIData ScienceDistributed SystemsEvent Driven ArchitectureGenerative AILangGraphLogistics & ProcurementMachine LearningNetworking InfrastructurePythonQuantizationSystem DesignTensorRTWeb DevelopmentvLLM

Frequently Asked Questions

How to apply for Applied AI Research Engineer: Benchmarking & Performance Economics at Akka?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

This position is open to freshers and experienced candidates.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Broken link or expired?
A

Akka

More jobs at Akka

Site Reliability Engineer

San Francisco, California

Forward Deployed Engineer (Australia)

San Francisco, California

Forward Deployed Engineer (N. America)

San Francisco, California

Share this Opening

Job Alerts for other

Receive email alerts whenever new other roles in San Francisco are posted.

Set Free Alert →

Similar Openings

Explore related active roles in other

View all
UrgentActively Hiring
Perfectserve
Practice Consultant I - US Remote
Perfectserve Verified
9+ years
₹692/mo
Remote
otherFull-timeRemote
Posted 17h ago
Apply Now
UrgentActively Hiring
P
Founding Engineer, Consumer Apps
Placerlabs Verified
6+ years
₹692/mo
United States, Remote
otherFull-timeRemote
Posted 17h ago
Apply Now
UrgentActively Hiring
S
Professional Services Architect
Sentinellabs Verified
5-8 years
Salary not disclosed
India
otherFull-time
Posted 17h ago
Apply Now

Applied AI Research Engineer: Benchmarking & Performance Economics

Akka · San Francisco

Apply on Company Website