You know that "just add RAG" is where the real work starts. You will own retrieval and context engineering across AI engagements, and you will measure your way to quality rather than guessing.
What you will do
- Design and tune retrieval pipelines: chunking, embeddings, hybrid search, re-ranking.
- Build the eval sets and metrics that turn "feels better" into a number.
- Engineer context and prompt strategies that hold up across edge cases.
- Manage the cost and latency budget of retrieval at production volume.
- Partner with data engineering on the pipelines that feed the index.
What we are looking for
- 6+ years engineering, with focused recent work on LLM and retrieval systems.
- Hands-on with vector stores (pgvector, Pinecone, Weaviate) and hybrid retrieval.
- Rigour with offline and online evaluation of retrieval and generation quality.
- Strong Python; comfort with the data plumbing behind an index.
- A scientific temperament – you trust measurements over vibes.
Nice to have
- Experience with fine-tuning or distillation where retrieval is not enough.
- Familiarity with eval tooling (Braintrust, Langfuse).
Apply for this role
Tell us about yourself and attach your resume. We review every application and reply within a few business days.