skip to content
Yonko Tsonev
researchexperimentstoolshownotesabout
sections
researchexperimentstoolshownotesabout

notes

Written down so it can be wrong.

The experiments run ahead of the writing, and mostly still do. What is here is what has been written up — including plans published before the results exist, which is the only way a prediction counts.

preregistrationAugust 2026

Affective Reinforcement Learning

Dual-policy optimization through a simulated internal state

The agent has two goals at once: do the task, and keep its own state up. That state is one number running from bad to good, and it also decides which of the two goals gets more say. Here is the plan, the predictions, and what would count as failure.

read →
← back to the research
Yonko Tsonevlearning · sentience · habit · conscienceindependent research · sofialast updated · august 2026privacygithub ↗linkedin ↗