Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There are almost certainly ways to fine-tune the model in ways that make it perform better on the Arena, but perform worse in other benchmarks or in practice. Usually that's not a good trade-off. What's being suggested here is that Meta is running such a fine-tuned version on the Arena (and reporting those numbers) while running models with different fine-tuning on other benchmarks (and reporting those numbers), while giving the appearance that those are actually the same models.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: