All jobs

Research Engineer - Agent Intelligence & Evaluation

Ixigo2w ago
New Delhi, DL, inOnsiteFull-time

Top focus

Research Scientist
  • We’re building self-healing voice agents for enterprise customer support within ixigo. The system has to know when it’s failing, why it’s failing
  • how to fix itself before a human notices. This fellowship sits at the intelligence layer behind that work. Voice agents fail in ways traditional software doesn't. ASR confidence drops on an accent and a tool call misfires. Latency breaks turn-taking and the LLM hallucinates a policy. A model swap silently regresses production and nobody catches it for a week. We're building self-healing voice agents for enterprise customer support. This role owns the intelligence layer: the evals that catch failures before shipping, the observability that traces them across the pipeline
  • the feedback loops that let agents fix themselves What you'll own
  • Evaluation infrastructure. Audio-native metrics for barge-in, prosody, and turn-taking. Adversarial datasets across accents and edge cases. LLM-as-judge rubrics for task success, tool-use correctness, and recovery.
  • Observability across the pipeline. Tracing that correlates audio, STT, LLM reasoning, tool calls, and TTS to a single conversation. Analysis and alerting that surfaces cascade failures instead of hiding them.
  • Self-improvement systems. Mine production traces for failure patterns, generate targeted training or prompt data, validate fixes with adversarial replay, and guardrail against regressions. Who we're looking for
  • 3 to 5 years in ML engineering, research engineering
  • applied research. Strong Python and modern ML tooling. Depth in at least two of: speech and audio models, LLM agent systems
  • eval or observability infrastructure.
  • You've shipped something non-trivial where research met production. You read papers, spot when a benchmark measures the wrong thing
  • translate ideas from Interspeech, ACL
  • NeurIPS into systems that run on real traffic. Publications welcome, not required. Nice to have Real-time systems or telephony experience. Work on RLHF, DPO
  • synthetic data pipelines. Familiarity with enterprise deployment (SOC 2, PII, data residency). What you'll get Senior seat on a small team where research and production aren't separate orgs. Real enterprise conversation data under proper governance. Meaningful equity, autonomy over tooling
  • support to publish. Candidates are responsible for safeguarding sensitive company data against unauthorized access, use
  • for reporting any suspected security incidents in line with the organization's ISMS (Information Security Management System) policies and procedures.

Required skills

PythonLLMSecurity
Posted on JobRush — the end-to-end AI job-search platform.