Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This is obviously the case if you apply much narrower criteria, most benchmarks of existing large datasets aren’t for character level tasks. That said, the synthetic noise section should be extremely interesting if not fully representative of your criteria.

Agree that tokenization isn’t a weakness for most general applications, disagree that it isn’t a weakness for the specific string manipulation task that the blog post is referencing



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: