LLMOps Engineer
About ALX Kenya
ALX Africa, a non-profit organisation under the ALX Foundation, is dedicated to unlocking the potential of Africa\'s digital future. Formerly part of Sand Tech Holdings, we\'ve embarked on an independent journey to provide world-class tech skills training and career acceleration programmes. Our mission is to bridge the digital divide, upskill and re-skill talent, and create a generation of innovative leaders. By 2030, we aim to empower 2 million Africans to secure sustainable tech careers.
Description
Qualifications
Specific Responsibilities
Evaluation Suites
Build and run eval suites per product, decomposed by failure mode, running on schedule and on every release — regression testing so nothing ships if it broke what worked.
Keep evals cost-effective as the product line grows.
Reporting, Data & Collaboration
Own the reporting loop — findings from evals and platform data in front of the team and stakeholders, including surfacing unintended or problematic model behaviour before learners do.
Steward the core datasets the team depends on, including classified customer-support data.
Partner with the AI Product Manager on instrumentation — they instrument the product, you build the evals over what is captured. This is a measurement role, not infrastructure — no model hosting or serving.
Skill Requirements – Essential
Python & data: solid Python and a data inclination, comfortable shaping and analysing messy LLM-generated data.
Decomposition: the ability to look at an AI product and decompose it into success and failure metrics.
Eval landscape: familiarity with Langfuse, RAGAS, DSPy, or similar — depth in one, awareness of the rest. These tools are learnable; we hire the fundamentals underneath them.
Desirable (not required): experience keeping evals cheap at scale; dashboarding and reporting; classical statistics.
Essential Traits for Success
You want to own a function, not execute tickets.
You communicate well and like collaborating, you will support every builder on the team.
You can point to any project, even a small one, where you measured an AI system honestly.