Senior Data Scientist Ãâ¢ã‚â€ã‚â” Clean Cooking (payg Lpg) At Sun King (formerly Greenlight Planet)
Sun King (Formerly Greenlight Planet)·Nairobi, kenya·Full Time·Internship
TechnologyFull TimeInternship
-
What Success Looks like
Your work will be measured against the commercial metrics of the PAYG LPG business, not model metrics alone. In your first 12 months you will be expected to:
Ship 2 - 3 data products into production use by commercial or operations teams, each with a measured impact on at least one of: activation rate, refill frequency, dormancy/churn, ARPU or cost-to-serve.
Establish the BU's approach to pricing and promotion analysis elasticity estimates, uplift measurement and experiment design that commercial leaders actually use to make decisions.
Build monitoring for the models you ship, with agreed retraining triggers and a clear owner for each.
What you will be expected to do:
Commercial
Develop first-hand understanding of customer and commercial needs across our markets, including time in the field with sales agents and customers.
Translate business problems into well-framed statistical questions, and present findings clearly to both technical and non-technical stakeholders up to BU leadership.
Prioritise ruthlessly: identify where a data product will move activation, refill frequency, retention or unit economics, and be willing to say where it won't.
Manage stakeholder expectations, product scope and delivery timelines for your own workstreams.
Technical
Design, build and evaluate machine learning models for business-critical use cases: churn/dormancy prediction, credit and payment-behaviour modelling, demand forecasting, anomaly detection and customer segmentation.
Apply probabilistic and Bayesian methods to quantify uncertainty and support decisions under uncertainty e.g. pricing elasticity, promotion uplift and media/marketing effectiveness.
Design and analyse experiments (A/B tests, geo tests, quasi-experiments) in field conditions where clean randomisation is often impossible.
Perform rigorous exploratory analysis, feature engineering and data wrangling on large structured and semi-structured datasets.
Partner with data and analytics engineering, who own pipelines and production infrastructure: you own the model from problem framing through validated, deployment-ready handoff, and jointly own monitoring once live.
Track and communicate model performance; identify degradation and recommend retraining or redesign.
Maintain clean, reproducible, well-documented code following team engineering standards.
What this role is not about:
Not a people-management role this is a senior IC position (a path to leading a small team may open as the function grows).
Not an MLOps/platform role you'll work to production standards, but pipeline and deployment infrastructure is owned by MLOps Engineer.
Not a reporting/BI role dashboarding exists in the analytics team; this role builds models and data products.
You might be a strong candidate if you have:
Degree in Computer Science, Statistics, Mathematics, Engineering, Economics or a closely related quantitative discipline. An advanced degree is a plus, not a requirement evidence of shipped impact matters more.
Commercial
A demonstrable track record of data products that measurably moved a business outcome you can walk us through the problem, the model, the decision it changed and the number it moved.
Ability to listen to and empathise with customers and colleagues, and to identify the P&L impact of a proposed data product before building it.
Strong communication and storytelling: you can carry a room of non-technical commercial leaders.
Experience managing upwards product needs, trade-offs and timelines.
Technical
5 - 8 years of hands-on experience in data science or applied ML roles, with at least 2 years owning data products end to end.
Strong command of classical ML (gradient boosting, regression, clustering, ranking, time-series forecasting) and the judgment to know when simple beats sophisticated.
Solid grounding in probabilistic modelling, Bayesian inference and uncertainty quantification, with working experience in a PPL such as PyMC or Stan.
High proficiency in Python (the standard scientific stack) and strong SQL, including complex multi-table queries and window functions.
Deep familiarity with model evaluation: cross-validation, calibration, and choosing business-aligned metrics over convenient ones.
Experience with experiment design and statistical hypothesis testing.
Comfortable working with cloud data warehouses and experiment-tracking tooling (we use AWS and MLflow; equivalents are fine).
Strongly preferred
Experience in PAYG, fintech lending, telco or other emerging-market consumer businesses you understand irregular incomes, mobile-money payment behaviour and thin, messy data.
Nice to have
Survival modelling, causal inference or marketing mix modelling (MMM).
Operations research / optimisation exposure (routing, scheduling) relevant to our last-mile delivery problems.
Familiarity with MLOps and model deployment on AWS (SageMaker, Lambda, ECS).