Engineering role

AI Engineer

Own the AI mechanics behind every feature we ship: retrieval, prompting, evaluation, guardrails and fallbacks. You make AI behave predictably in front of real users, and you can prove it.

  • Senior
  • Full-time
  • Remote
About the role

What this job actually is.

Most AI features fail quietly. They answer confidently and wrongly, they cost more than anyone expected, or nobody can tell whether they work. Your job is to make sure ours do not.

You will design the retrieval, prompting and evaluation behind the features we build for clients, and you will define what good looks like before anything goes live. You will also be the person who says a feature should not ship yet, and shows why.

You work closely with our designers and full-stack engineers, so the model's behavior matches the experience on the screen.

What you'll do

The work you will own.

A typical engagement asks for all of this. Some weeks lean harder on one part than another.

  1. 01Design retrieval, prompting and tool use pipelines for features running inside client products.
  2. 02Build evaluation sets and pipelines, and set the quality bar a feature must pass before release.
  3. 03Add guardrails, fallbacks and human review steps where a wrong answer would cost the user something real.
  4. 04Pick the smallest model and simplest approach that does the job, and track cost and latency from day one.
  5. 05Monitor features in production and investigate failures with logs, traces and real user examples.
  6. 06Advise clients plainly on which AI ideas are worth building and which are not.
  7. 07Work with design so confidence, sources and uncertainty are shown honestly in the interface.
  8. 08Document how each system works so the client's team can maintain and improve it.
What you bring

The bar for this role.

Don't meet every line? Apply anyway. We hire for judgment over a checklist.

Must have

  • 4 or more years building production software, with recent hands-on work on LLM applications.
  • Deep practical experience with model APIs, prompting, retrieval and structured outputs.
  • You have built evaluation pipelines. You do not ship AI you cannot measure.
  • Strong engineering fundamentals in Python or TypeScript, including APIs, testing and observability.
  • Pragmatic about models and tools, with a preference for the boring option that works.
  • Comfortable saying this should not ship, and backing it with data.

Nice to have

  • Experience tuning retrieval with pgvector, Pinecone or a similar vector store.
  • Hands-on work with agent and tool use workflows in production.
  • Background in search, ranking, recommendations or classic ML systems.
  • You have run an AI feature at scale and can talk about the failures you saw.
Tools you'll use
  • Python
  • TypeScript
  • OpenAI and Anthropic APIs
  • pgvector
  • Eval tooling
  • Tracing and observability
Your first 90 days

How you will ramp up.

A clear path from joining to leading, with a senior partner beside you the whole way.

  1. First 30 days

    Learn our evaluation habits

    Review the AI systems on active engagements, learn how we evaluate and release, and improve one live feature with measured results.

  2. By day 60

    Own a system

    Design and ship the retrieval and evaluation for one client feature, with the release bar agreed before launch.

  3. By day 90

    Set the standard

    Own the AI quality bar across engagements and improve the shared tooling the whole team uses to test and monitor.

How to apply

Tell us what you've shipped.

If you're interested in joining our team, send a message to careers@uplayer.agency.

What to include
  • The role you are applying for, or the role you think we are missing.
  • Links to work you shipped, and what your part in it was.
  • A few lines on a time you decided not to build something, and why that was the right call.

We read every application.

Ready to apply?

AI Engineer

Senior, full-time, Remote, GMT-6 to GMT+6.

Apply for this role