Efficient Benchmarking Is Just Feature Selection and Multiple Regression
I'm a Member of Technical Staff at Cohere, working on pretraining evals. Previously, I spent four years on a PhD in probabilistic machine learning at the University of Bristol's Compass CDT (viva pending!). My research interests include probabilistic modelling, deep learning, and using traditional statistics to improve LLM evaluation procedures.
Feel free to get in touch!
Learning Generation Orders for Masked Discrete Diffusion Models via Variational Inference
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
Special Report: Compass Away Day 2024
Sam Bowyer // Bristol