AI Data Researcher - Python, LLM, RAG
Software Engineering, Data Science
Barbados · Mexico · Dominica · Dominican Republic · Haiti · Jamaica · South America · Central America · Cuba · Antigua and Barbuda · The Bahamas · Belize · Guyana · Grenada · St Kitts & Nevis · St Vincent and the Grenadines · Suriname · St Lucia · Trinidad and Tobago · Remote
Posted on Sep 16, 2026
Requirements
Must-haves
- 4+ years of data research, applied analytics, or evaluation experience
- Experience with LLMs, prompting, RAG, and retrieval concepts
- Proficiency with Python
- Proficiency with SQL
- Experience in maintaining high-quality datasets or evaluation programs
- Ability to navigate ambiguous problems independently
- Strong communication skills in both spoken and written English
What makes you stand out
Fully own and drive the quality of the evaluation system
Nice-to-haves
- Startup experience
- Experience with LLM-as-a-judge evaluation
- Familiarity with (LangSmith, Ragas, DeepEval, Braintrust, Evidently, etc.)
- Knowledge of embeddings, search, reranking, or hybrid retrieval
- Experience creating evaluation rubrics, dashboards, or lightweight automation
- Bachelor's Degree in Computer Engineering, Computer Science, or equivalent
What you will work on
- Own the quality of the LLM and Retrieval-Augmented Generation (RAG) evaluation system
- Run structured LLM and RAG evaluations and perform root-cause analysis
- Distinguish between retrieval, ranking, source, prompt, model, and citation failures
- Maintain evaluation datasets, including quality reviews, versioning, coverage, and new production cases
- Refine prompts, judge rubrics, and LLM-as-a-judge workflows
- Design controlled experiments with clear baselines, metrics, and success criteria
- Investigate practical solutions to recurring RAG challenges
- Automate repetitive analysis (e.g., Python, SQL, notebooks, APIs)
- Communicate findings and recommended actions to Product, Engineering, and leadership