Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Continuous RL in a sense. There maybe an undiscovered additional scaling law around models doing what you describe; continuous LLM-as-self-judge, if you will.

Provided it can be determined why a user ended the chat, which may turn out to be possible in some subset of conversations.



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: