Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
AI on Trial: Fine-Tuning LLMs to Judge other LLMs; data, results and challenges
We will present data, experiments, results, and challenges of training language models to evaluate other models, plus insights on upcoming experiments.
We’re sharing the progress of our work on training language models to evaluate other language models. Our session begins with an overview of the training data, then moves to detailed discussions of our initial experiments, their outcomes, and what we’ve learned from them. We’ll address the challenges encountered and offer a glimpse into our forthcoming experiments. The aim is to foster a transparent exchange on the practicalities and potential of evaluator models in AI.
Atla provides agent observability, error detection, trace analysis, and performance experimentation.
Fine-tuning LLMs with preference data to act as robust LLM evaluators.
Compose Email
Loading recent emails...