Continuous RL in a sense. There maybe an undiscovered additional scaling law around models doing what you describe; continuous LLM-as-self-judge, if you will.
Provided it can be determined why a user ended the chat, which may turn out to be possible in some subset of conversations.
Provided it can be determined why a user ended the chat, which may turn out to be possible in some subset of conversations.