Ops & Evaluation · 32 / 48
Benchmarking Models
Goal: Learn to benchmark models on standard tasks like MMLU and HumanEval so you can pick the right model for a job with data instead of hype.
Do:
Read this guide: https://www.ibm.com/think/topics/llm-benchmarks
Run your first benchmarks with this tool https://github.com/EleutherAI/lm-evaluation-harness
Saved in this browser.