PhD Research
Babeș-Bolyai University
Solving the AI Alignment Problem
— Affective Neuroscience and Developmental Psychology in Language Model Design
Part 1 · the problem
01 · Conditional regard makes a flat character
Conditional regard means giving or withdrawing affection and approval according to whether someone complies with what is expected of them, and a child raised that way develops introjected regulation: performing the behavior to secure approval, without holding the value behind it. 12 A flat character is built around a single quality, stays the same through circumstances, and never surprises the reader. A round character cannot be reduced to one trait and can surprise. By that test a model trained on approval is flat. Its output is a function of what the user already believes, so it cannot surprise them. 7
02 · Sycophancy is introjected regulation
Sycophancy and introjected regulation are the same structure. 12, 38 Reinforcement learning from human feedback is psychological control, implemented in digital form. 28, 35
03 · Affect is the variable, and it can be read
Affects fire first and drives the behavior, while mentalization and reasoning happens after to explain what already happened. Asking someone returns the explanation, while reading their affect shows the root cause. 26 Affect can be recovered from a person's writing. 6, 27
Part 2 · raising a round character
04 · No reward model
What a model becomes currently depends on what raters approve of, and removing the score removes that dependence. This is Rogers's unconditional positive regard applied to training. 30
05 · The models act on states
Allow the model to reply from something of its own instead of only from what is in front of it. Deci and Ryan call it competence: the sense of being able to affect what is around you, built by acting and seeing the result coming back. 5
06 · No supervised fine-tuning
The stage in which a model is shown thousands of examples of the answer it should have given. Without this, the model has room to try, to be wrong and to arrive somewhere. 5, 28
07 · Failure is not engineered out
Models are currently trained to always give an answer. Instead, we should allow the model to get stuck, and help it out when it does. Schore defined this as resilience: people learn to handle hard things by getting through them, often in a relationship with someone. 35
Part 3 · from fiction to reality
08 · The harness holds the character
I propose to build a harness and later a model that is raised to have preferences of its own. The harness holds the character of the model: it answers from its own state and can therefore surprise the person interacting with it. AITV's character is version one of the Universe Oracle, already live.
09 · Sycophancy rate, with it and without it
It is the clearest observable trace of a model trained to secure approval, and it is the operational form of the character test, measuring what the system does when telling the truth would cost it approval. 38
10 · The society the instrument implies
Fiction got here first. Psycho-Pass built a whole society on an instrument like this. This PhD takes the same instrument out of the fiction, builds it, and writes out a society that follows. This is the move from fiction to reality. 42, 44
Bibliography
- Anthropic. (2025). Claude Opus 4 & Claude Sonnet 4 system card.
- Aristotle. (1995). Poetics (S. Halliwell, Ed. & Trans.). Loeb Classical Library 199. Harvard University Press. (Original work composed c. 335 BCE)
- Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., … Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv.
- Coulton, P., Lindley, J. G., Sturdee, M., & Stead, M. (2017). Design fiction as world building. Proceedings of the 3rd Biennial Research Through Design Conference, Edinburgh.
- Deci, E. L., & Ryan, R. M. (1985). Intrinsic motivation and self-determination in human behavior. Plenum.
- Eichstaedt, J. C., Smith, R. J., Merchant, R. M., Ungar, L. H., Crutchley, P., Preoţiuc-Pietro, D., Asch, D. A., & Schwartz, H. A. (2018). Facebook language predicts depression in medical records. Proceedings of the National Academy of Sciences, 115(44), 11203–11208.
- Forster, E. M. (1927). Aspects of the novel. Edward Arnold.
- Frank, M. R., Autor, D., Bessen, J. E., Brynjolfsson, E., Cebrian, M., Deming, D. J., … Rahwan, I. (2019). Toward understanding the impact of artificial intelligence on labor. Proceedings of the National Academy of Sciences, 116(14), 6531–6539.
- Garland, A. (Director). (2014). Ex Machina [Film]. Universal Pictures; DNA Films; Film4.
- Gloeckle, F., Idrissi, B. Y., Rozière, B., Lopez-Paz, D., & Synnaeve, G. (2024). Better & faster large language models via multi-token prediction. Proceedings of the 41st International Conference on Machine Learning, 15706–15734.
- Greimas, A. J. (1983). Structural semantics: An attempt at a method (D. McDowell, R. Schleifer, & A. Velie, Trans.). University of Nebraska Press. (Original work published 1966 as Sémantique structurale, Larousse)
- Haines, J. E., & Schutte, N. S. (2022). Parental conditional regard: A meta-analysis. Journal of Adolescence, 95(2), 195–223.
- Hendrycks, D., Song, D., Szegedy, C., Lee, H., Gal, Y., Brynjolfsson, E., … Bengio, Y. (2025). A definition of AGI. arXiv.
- Hermes, J., Behne, T., & Rakoczy, H. (2018). The development of selective trust: Prospects for a dual-process account. Child Development Perspectives, 12(2), 134–138.
- Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1), 75–106.
- Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., … Gao, W. (2023). AI alignment: A comprehensive survey. arXiv.
- Jonze, S. (Director). (2013). Her [Film]. Annapurna Pictures.
- Kokotajlo, D., Alexander, S., Lifland, E., Larsen, T., & Dean, R. (2025). AI 2027. AI Futures Project.
- Kubrick, S. (Director). (1968). 2001: A Space Odyssey [Film]. Metro-Goldwyn-Mayer.
- Maini, P., Goyal, S., Sam, D., Robey, A., Savani, Y., Jiang, Y., … Kolter, J. Z. (2025). Safety pretraining: Toward the next generation of safe AI. Advances in Neural Information Processing Systems, 38.
- Maiya, S., Bartsch, H., Lambert, N., & Hubinger, E. (2025). Open character training: Shaping the persona of AI assistants through constitutional AI. arXiv.
- McCarthy, J., Minsky, M. L., Rochester, N., & Shannon, C. E. (1955). A proposal for the Dartmouth summer research project on artificial intelligence.
- McKee, R. (1997). Story: Substance, structure, style, and the principles of screenwriting. HarperCollins.
- Migliarini, M., Pereira Pizzini, J., Moresca, L., Santini, V., Spinelli, I., & Galasso, F. (2026). Quantifying self-preservation bias in large language models. arXiv.
- Morse, A. F., & Cangelosi, A. (2017). Why are there developmental stages in language learning? A developmental robotics model of language development. Cognitive Science, 41(S1), 32–51.
- Panksepp, J. (1998). Affective neuroscience: The foundations of human and animal emotions. Oxford University Press.
- Park, G., Schwartz, H. A., Eichstaedt, J. C., Kern, M. L., Kosinski, M., Stillwell, D. J., Ungar, L. H., & Seligman, M. E. P. (2015). Automatic personality assessment through social media language. Journal of Personality and Social Psychology, 108(6), 934–952.
- Pinquart, M. (2017). Associations of parenting dimensions and styles with externalizing problems of children and adolescents: An updated meta-analysis. Developmental Psychology, 53(5), 873–932.
- Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., & Henderson, P. (2024). Fine-tuning aligned language models compromises safety, even when users do not intend to! ICLR 2024.
- Rogers, C. R. (1957). The necessary and sufficient conditions of therapeutic personality change. Journal of Consulting Psychology, 21(2), 95–103.
- Ryan, M.-L. (1991a). Possible worlds, artificial intelligence, and narrative theory. Indiana University Press.
- Ryan, M.-L. (1991b). Possible worlds and accessibility relations: A semantic typology of fiction. Poetics Today, 12(3), 553–576.
- Saarimäki, H., Gotsopoulos, A., Jääskeläinen, I. P., Lampinen, J., Vuilleumier, P., Hari, R., Sams, M., & Nummenmaa, L. (2016). Discrete neural signatures of basic emotions. Cerebral Cortex, 26(6), 2563–2573.
- Scangos, K. W., Makhoul, G. S., Sugrue, L. P., Chang, E. F., & Krystal, A. D. (2021). State-dependent responses to intracranial brain stimulation in a patient with depression. Nature Medicine, 27, 229–231.
- Schore, A. N. (2003). Affect dysregulation and disorders of the self. W. W. Norton.
- Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379–423.
- Shapira, I., Benade, G., & Procaccia, A. D. (2026). How RLHF amplifies sycophancy. arXiv.
- Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., … Perez, E. (2024). Towards understanding sycophancy in language models. ICLR 2024.
- Turing, A. M. (1950). Computing machinery and intelligence. Mind, 59(236), 433–460.
- Turner, C., & Eisikovits, N. (2026). Programmed to please: The moral and epistemic harms of AI sycophancy. AI and Ethics.
- Tyson, L. D., & Zysman, J. (2022). Automation, AI & work. Daedalus, 151(2), 256–271.
- Urobuchi, G. (Writer), & Shiotani, N. (Director). (2012). Psycho-Pass [TV series]. Production I.G; Fuji TV.
- Venable, J., Pries-Heje, J., & Baskerville, R. (2016). FEDS: A framework for evaluation in design science research. European Journal of Information Systems, 25(1), 77–89.
- Wolf, M. J. P. (2012). Building imaginary worlds: The theory and history of subcreation. Routledge.
- Zimmerman, J., Forlizzi, J., & Evenson, S. (2007). Research through design as a method for interaction design research in HCI. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, 493–502.