name

Senior AI Backend Engineer - Agent Evaluation & Quality

Remote, United States remote Entry Salary not listed
remote Technology & IT Curated
Sign in to apply Free account — we bring you straight back to this role.

About the role

About the role
We run production multi-agent systems that handle real work for a large base of users. As those systems grow, our biggest constraint is confidence: we need to know how well the agents perform, catch regressions before they ship, and keep quality steady as we release. This role owns that.
You'll build the evaluation systems behind our agents - the judges, test harnesses, and simulators that tell us whether an agent is working and where it's failing. The goal is to let us ship agents faster because we can trust what the evaluation tells us.
Evaluation is the focus, but it won't be the boundary. Because you'll understand the agents' failure modes better than anyone there will also be opportunities to contribute to agent development itself, building and improving the agents alongside the systems that evaluate them.
Responsibilities
Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode.
Make the release gate real. Build per-PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes.
Build user simulators to generate test coverage and adversarial cases before real users hit them.
Turn production signal into improvement - pipe real failures back into evaluation sets so the system compounds over time.
Partner with product to turn "what good looks like" into concrete, measurable criteria.
Grow into agent development - contribute to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.
Requirements
Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering.
Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break.
A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it.
Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability.
5+ years software engineering, with recent hands-on LLM/agent work.
Nice to have
Direct experience evaluating LLM/agent systems - offline/online eval, LLM-as-judge, systematic regression testing.
Observability tooling (Arize, LangSmith, or similar).
Arabic language / NLP experience.
E-commerce or merchant-facing product experience.
Originally posted on Himalayas

Interview prep

Walk in with sharper answers.

Use this as a quick practice sheet before you speak with the employer.

Role
Technology & IT API integration Javascript Python Senior Backend remote

Likely questions

  1. Tell us about work you have done that is close to the Senior AI Backend Engineer - Agent Evaluation & Quality role.
  2. How would you approach your first 30 days at name?
  3. Which of API integration, Javascript and Python have you used recently, and what did it help you achieve?
  4. Describe a time you solved a problem without waiting to be told exactly what to do.
  5. How do you stay organised and communicate clearly when working remotely?

Prepare before the call

  • A recent example that proves your experience with API integration, Javascript and Python.
  • One short story with a problem, your action, and the result.
  • Two examples that show the strengths listed on your CV.
  • A clear reason why this role and company interest you.
  • Your availability, preferred work style, and salary expectations.

Ask them

  • What would success look like in the first 90 days?
  • What are the main problems this hire should help solve?
  • How does the team give feedback and measure good work?
  • What does a normal working week look like for this role?
Practice line

I am interested in the Senior AI Backend Engineer - Agent Evaluation & Quality role because I can bring practical experience in API integration, Javascript and Python, learn the team quickly, and contribute to the outcomes name needs from this hire.

Related jobs.

More roles from this company or category.