AI Data Researcher - Python, LLM, RAG

Strider
Strider

Software Engineering, Data Science

Barbados · Mexico · Dominica · Dominican Republic · Haiti · Jamaica · South America · Central America · Cuba · Antigua and Barbuda · The Bahamas · Belize · Guyana · Grenada · St Kitts & Nevis · St Vincent and the Grenadines · Suriname · St Lucia · Trinidad and Tobago · Remote

Posted on Sep 16, 2026

Requirements

Must-haves

  • 4+ years of data research, applied analytics, or evaluation experience
  • Experience with LLMs, prompting, RAG, and retrieval concepts
  • Proficiency with Python
  • Proficiency with SQL
  • Experience in maintaining high-quality datasets or evaluation programs
  • Ability to navigate ambiguous problems independently
  • Strong communication skills in both spoken and written English

What makes you stand out

Fully own and drive the quality of the evaluation system

Nice-to-haves

  • Startup experience
  • Experience with LLM-as-a-judge evaluation
  • Familiarity with (LangSmith, Ragas, DeepEval, Braintrust, Evidently, etc.)
  • Knowledge of embeddings, search, reranking, or hybrid retrieval
  • Experience creating evaluation rubrics, dashboards, or lightweight automation
  • Bachelor's Degree in Computer Engineering, Computer Science, or equivalent

What you will work on

  • Own the quality of the LLM and Retrieval-Augmented Generation (RAG) evaluation system
  • Run structured LLM and RAG evaluations and perform root-cause analysis
  • Distinguish between retrieval, ranking, source, prompt, model, and citation failures
  • Maintain evaluation datasets, including quality reviews, versioning, coverage, and new production cases
  • Refine prompts, judge rubrics, and LLM-as-a-judge workflows
  • Design controlled experiments with clear baselines, metrics, and success criteria
  • Investigate practical solutions to recurring RAG challenges
  • Automate repetitive analysis (e.g., Python, SQL, notebooks, APIs)
  • Communicate findings and recommended actions to Product, Engineering, and leadership