Nikola Milosevic

Nikola Milosevic

Ph.D. Candidate, MPI CBS, Leipzig

I’m a Ph.D. researcher at MPI CBS working at the intersection of reinforcement learning, control, and optimization, with applications in theoretical neuroscience, AI safety, and robotics.

My work focuses on the mathematical foundations of reinforcement learning, in particular the geometry of policy optimization, and on turning them into algorithms with guarantees: agents that satisfy safety constraints throughout training, objectives beyond reward maximization, and principled models of adaptive behavior such as active inference.

Selected publications

Active Inference as a Convex Markov Decision Process

Nikola Milosevic, Nicolás Hinrichs, Nico Scherf

International Workshop on Active Inference (IWAI 2026)

Abstract

Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle. We frame AIF as policy optimization and show that, for closed-loop control policies, EFE minimization can be formulated as a convex Markov decision process (MDP). This perspective reveals that policy-dependent reward prediction errors transmit natural gradients of the expected free energy backwards in time rather than up a hierarchy. Finally, we show that coupling world-model learning with policy optimization gives active inference the structure of performative reinforcement learning. Together this places EFE minimization within modern reinforcement learning and optimization theory and opens a route toward principled algorithms for active inference.

Physical embodiment enables information processing beyond explicit flow sensing in active matter

Diptabrata Paul, Nikola Milosevic, Nico Scherf, Frank Cichos

Science Advances, 12(11), eaec0783

Abstract

We show that physical embodiment in active matter systems enables information processing capabilities that exceed what is possible through explicit flow sensing alone. Using microswimmers as a model system, we demonstrate that the body itself acts as a computational resource, coupling sensory and motor degrees of freedom in ways that simplify the control problem.

Embedding Safety into RL: A New Take on Trust Region Methods

Nikola Milosevic, Johannes Müller, Nico Scherf

International Conference on Machine Learning (ICML), PMLR 267:44199-44224

Abstract

Reinforcement Learning (RL) agents can solve diverse tasks but often exhibit unsafe behavior. Constrained Markov Decision Processes (CMDPs) address this by enforcing safety constraints, yet existing methods either sacrifice reward maximization or allow unsafe training. We introduce Constrained Trust Region Policy Optimization (C-TRPO), which reshapes the policy space geometry to ensure trust regions contain only safe policies, guaranteeing constraint satisfaction throughout training. We analyze its theoretical properties and connections to TRPO, Natural Policy Gradient (NPG), and Constrained Policy Optimization (CPO). Experiments show that C-TRPO reduces constraint violations while maintaining competitive returns.

Talks & posters

The Duality of Perception and Action in Brains, Minds, and Machines UpcomingNov 2026Talk · Reinforcement Learning Coffee, PLUS Salzburg
Active Inference as a Convex Markov Decision Process UpcomingOct 2026International Workshop on Active Inference (IWAI 2026) · Madrid, Spain
The Geometry of Nonlinear Reinforcement LearningNov 2025Lightning talk · Workshop on Geometry, Topology, and Machine Learning (GTML 2025), Leipzig
On the Generality of Relative Entropy Policy IterationOct 2025Talk · MiS/ScaDS/CBS Math and AI Meeting, Leipzig University
Embedding Safety into RL: A New Take on Trust Region MethodsSep 2025European Workshop on Reinforcement Learning (EWRL 2025) · Tübingen, Germany
Embedding Safety into RL: A New Take on Trust Region MethodsJul 2025Poster · ICML 2025 · Vancouver, Canada
Central Path Proximal Policy OptimizationJul 2025Poster · Exploration in AI Today Workshop, ICML 2025 · Vancouver, Canada

Teaching & supervision

Teaching Assistant, Advanced Deep LearningpresentLeipzig University · Lectures on deep reinforcement learning and deep latent variable models for Master’s students
Teaching Assistant, Systems Theory2020 — 2021HTWK Leipzig · Tutorials and exam preparation for third-semester students