CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Back to all jobs
  1. Home
  2. Jobs
  3. Senior AI Engineer — Inference & Agent Systems
AA
Arcana Analytics

Senior AI Engineer — Inference & Agent Systems

United States
₹10.4L – ₹17.3L/mo
Full-time
Posted 5d ago
2 views
Actively Hiring Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Senior AI Engineer — Inference & Agent Systems role at Arcana Analytics.

Apply for this position

Apply on Company Website Email: [email protected]
Notice a broken link or wrong info?

Job Description

  • Title: Applied AI Engineer — Inference & Agent Systems

  • Location:
    United States

What We're Building

Arcana is building AI agents that synthesize information across heterogeneous sources and deliver structured, reasoned answers in real time. The product works if the agents are fast, reliable, and correct, not approximately correct

  • Our stack: Go + Temporal for orchestration, a Plan-Execute-Synthesize agent architecture, and an evaluation harness we use to measure every regression. The problems are hard. The latency bar is aggressive. The accuracy requirements are unforgiving.

The Work

Inference Optimization

- Drive TTFT below 400ms for multi-step agent pipelines

- Streaming optimization: first token to user while sub-agents are still running

- KV cache strategy, prompt compression, dynamic context window management

- Multi-provider routing: model selection by latency, cost, and task type across OpenAI, Anthropic, Gemini, and open-weight models

Agent Architecture

- Design and implement Plan-Execute-Synthesize pipelines that run sub-agents in parallel DAGs, not sequential chains

- Build reliable orchestration on top of Temporal: retries, timeouts, partial failure recovery, idempotency

- Structured output enforcement: JSON schema validation, retry loops on malformed LLM output, graceful degradation

- Tool call design: schema design that LLMs actually follow reliably across providers

Evaluation & Harness

- Own the eval framework end to end: ground truth datasets, automated scoring pipelines, regression detection on every PR

- LLM-as-judge pipelines for qualitative output assessment

- Latency regression testing - p50/p95/p99 tracked across every deployment

- Adversarial test case design: ambiguous queries, missing data, conflicting sources, malformed tool responses

Infrastructure

- Model serving and cold start optimization

- Async worker architecture for parallel sub-agent execution

- Observability: trace every token, every tool call, every synthesis step

What We're Looking For

You've built something that runs in production at a meaningful scale and you understand why it's fast (or why it isn't).

  • Strong signal:

- You've worked on inference pipelines where TTFT was the primary metric and you moved it meaningfully

- You've built multi-step agent systems and you know where they break not from reading papers but from watching them fail in production

- You've written eval harnesses from scratch and you have opinions about what makes a ground truth dataset actually useful

- You've debugged LLM non-determinism in production and built systems resilient to it

- You've worked with streaming LLM responses and built infrastructure around partial output handling

  • Weaker signal (but not disqualifying):

- You've fine-tuned models but haven't shipped inference systems

- You've used LangChain/LlamaIndex but haven't built the layer underneath

- Strong ML research background without systems exposure

Stack familiarity

(we care more about depth than match): Go, Python, Temporal, Kafka, PostgreSQL, Docker

Why This Role

The problems here don't have blog posts about them yet. Parallel agent DAG execution under hard latency budgets, streaming synthesis across partial sub-agent results, eval harnesses for non-deterministic multi-step systems: these are genuinely unsolved at production quality. Small team. High ownership. Every engineer's decisions ship to production.

Who We Want to Hear From

You've shipped inference systems at:

- A real-time AI product (search, coding assistant, chat at scale)

- A model serving infrastructure company

- An agent platform (any domain)

Or you've built eval/harness infrastructure that a team of 10+ engineers actually trusted to catch regressions.

Apply

Send to: [[email protected]]

  • Include:
  1. One system you built where latency was the primary constraint what you measured, what you changed, what moved
  2. Link to anything public (code, writing, talks)
  3. No cover letter required

We respond to every application.

Frequently Asked Questions

How to apply for Senior AI Engineer — Inference & Agent Systems at Arcana Analytics?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

The salary for this role is USD 150,000 - 250,000/yr per annum.

What experience is required?

This position is open to freshers and experienced candidates.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Send Email Application
Broken link or expired?
AA

Arcana Analytics

More jobs at Arcana Analytics

Data Associate

Coimbatore

Portfolio Analytics Specialist

London Area, United Kingdom (Remote)

Data Analyst

Coimbatore/Bangalore/Remote

Share this Opening

Job Alerts for ai_engineering

Receive email alerts whenever new ai_engineering roles in United States are posted.

Set Free Alert →

Similar Openings

Explore related active roles in ai_engineering

View all
UrgentActively Hiring
Finc
Senior AI Engineer
Finc Verified
0-2 Yrs
Salary not disclosed
Bengaluru
ai_engineeringFull-time
Posted 22h ago
Apply Now
UrgentActively Hiring
Fa-espx-saasfaprod1
Senior AI Engineer
Fa-espx-saasfaprod1 Verified
8+ years
Salary not disclosed
Pune, Maharashtra, India
ai_engineeringFull-time
Posted 22h ago
Apply Now
UrgentActively Hiring
Grafana Labs
Staff AI Engineer - Grafana AI/ML | Canada | Remote
Grafana Labs Verified
0-2 Yrs
Salary not disclosed
Canada (Remote)
ai_engineeringFull-timeRemote
Posted 22h ago
Apply Now

Senior AI Engineer — Inference & Agent Systems

Arcana Analytics · United States

Apply on Company Website