
Hi 👋 I am a Lead Research Scientist at Thomson Reuters Foundational Research, where I am the research lead for LLM Evaluations as part of the training of Thomson-1, Thomson Reuters’ own frontier-class foundation model (~400B parameters).
Most of my work now sits on the evaluation and verification of LLMs and agents — how do we actually know whether a model works? In practice that means decomposing what a good answer means in high-stakes domains, grounding model claims in authoritative evidence, and making rigorous evaluation orders of magnitude cheaper to run 🚀. I also work on how LLMs reason and search over hard problems, and on who they end up aligning to when demands compete. Throughout, I come at this from a data-centric angle: what you train and evaluate on determines what you can actually trust 🦾.
I completed my PhD in Machine Learning at the University of Cambridge, supervised by Prof. Mihaela van der Schaar. I hold a Masters degree from Cornell University working on Bayesian Deep Learning, as well as a Masters from the University of the Witwatersrand (South Africa) working on Signal Processing & ML for Parkinson’s Disease. I also hold a dual-bachelors in Information Engineering & Biomedical Engineering from the University of the Witwatersrand (South Africa).
Previously, my industry experience includes time as an ML Researcher at AstraZeneca working on LLM verification, test-time scaling and uncertainty quantification. Before my PhD, I worked on production ML systems as a Data Scientist working on Computer Vision at Shutterstock (USA) and as an ML Engineer working on NLP at Multichoice (Africa’s largest multimedia company).
Outside the research itself, I was named one of the Mail & Guardian’s Top 200 Young South Africans 🇿🇦
Always happy to chat about research — feel free to reach out!
Research interests
- LLM Evaluation
- LLM Agents & Verification
- LLM Reasoning
- Data-Centric AI
- Responsible AI & Alignment
- Uncertainty Quantification
- Synthetic Data
Publications
Please find some of my publications below (a more up-to-date list can be found on google scholar). “*” denotes equal contribution.
LLM Evaluation & Verification
5
Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification
To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands
Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings
Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models
LLM Agents, Reasoning & Programs
6
Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces
Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback
Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema Matching
Position: What's the next frontier for Data-centric AI? Data Savvy Agents!
Matchmaker: Self-Improving Large Language Model Programs for Schema Matching
Large Language Models to Enhance Bayesian Optimization
Data-Centric AI
11
Towards Human-Guided, Data-Centric LLM Co-Pilots
Going Beyond Static: Understanding Shifts with Time-Series Attribution
Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World Environments
You can't handle the (dirty) truth: Data-Centric Insights Improve Pseudo-Labeling
Dissecting Sample Hardness: A Fine-Grained Analysis of Hardness Characterization Methods for Data-Centric AI
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
DAGnosis: Localized Identification of Data Inconsistencies using Structures
Navigating Data-Centric Artificial Intelligence with DC-Check: Advances, Challenges, and Opportunities
TRIAGE: Characterizing and auditing training data for improved regression
Data-IQ: Characterizing subgroups with heterogeneous outcomes in tabular data
Data-SUITE: Data-centric identification of in-distribution incongruous examples
Uncertainty Quantification
8
Active Learning with LLMs for Partially Observed and Cost-Aware Scenarios
Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise
U-PASS: An uncertainty-guided deep learning pipeline for automated sleep staging
Automated remote sleep monitoring needs uncertainty quantification
What is Flagged in Uncertainty Quantification? Latent Density Models for Uncertainty Categorization
Improving Adaptive Conformal Prediction Using Self-Supervised Learning
MCU-Net: A framework towards uncertainty representations for decision support system patient referrals in healthcare contexts
Towards calibrated and scalable uncertainty representations for neural networks
Synthetic Data
3
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in ultra low-data regimes
Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test Data
Reimagining Synthetic Tabular Data Generation through Data-Centric AI: A Comprehensive Benchmark
Healthcare & Responsible AI
4
Tables2Traces: Distilling Tabular Data to Improve LLM Reasoning in Healthcare
Unlocking Historical Clinical Trial Data with ALIGN: A Compositional Large Language Model System for Medical Coding
Generalization—a key challenge for responsible AI in patient-facing clinical applications
Modeling Disagreement in Automatic Data Labelling for Semi-Supervised Learning in Clinical Natural Language Processing
Causal Inference & Structure Learning
2
Differentiable and transportable structure learning
Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations
Biomedical Signal Processing & Applied ML
9
PEMS: Custom Neural Machine Translation System—Making subtitling of Portuguese TV shows and movies on the African continent work
Machine learning discrimination of Parkinson's Disease stages from walker-mounted sensors data
Automated machine vision enabled detection of movement disorders from hand drawn spirals
Automated and interpretable m-health discrimination of vocal cord pathology enabled by machine learning
Automated stage discrimination of Parkinson's disease
A comparison of footfall detection algorithms from walker mounted sensors data
Feasibility of an instrumented walker to quantify treatment effects on Parkinson's patient gait
Custom force sensor and sensory feedback system to enable grip control of a robotic prosthetic hand
Quadcopter Control using Intelligent Control Methods
Tutorials
3
Clinical AI in the real-world: From Data-centric AI to Dynamic Learning
Data-Centric AI for reliable and responsible AI: from theory to practice
Data-Centric AI: Foundation, Frontiers and Applications in the quest for robust and reliable AI systems
Thesis
1
Data-Centric AI for Reliable and Trustworthy Machine Learning
Awards & honours
- Best Paper Award ICLR 2026 AIWILD Workshop — Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification 2026
- Top 200 Young South Africans Mail & Guardian 2023
- Best Research Poster Presentation Future of Data-Centric AI Conference 2023
News
Show 26 earlier items Hide earlier items
Featured talk
Watch on SlidesLiveInvited talks
- GE HealthcareSynthetic data Oct 2024
- Nature Digital MedicineGeneralization as a key challenge for Responsible AI June 2024
- StanfordAn uncertainty estimation lens on Data-centric AI June 2024
- ICTP Advanced ML summer schoolData-Centric AI May 2024
- AppleData-Centric AI — Data characterization & Synthetic Data April 2024
- Microsoft Research CambridgeData-Centric AI November 2023
- Future of Data-Centric AI Conference TalkData-IQ June 2023
- Discovery Limited Invited TalkData-Centric AI Feb 2023
- AstraZeneca AI Journal Club Invited TalkData-Centric AI Nov 2022
- Queen Mary University London CogSci Invited TalkData-Centric AI Oct 2022
- University of Cape Town Invited TalkData-Centric AI Oct 2022