keyword
projection algorithm
The projection algorithm is an iterative apprenticeship learning method designed to learn an agent policy that matches the performance of an expert demonstrator in a Markov decision process without requiring prior knowledge of the underlying reward function. Operating under the assumption that the unknown reward is expressible as a linear combination of known state features, the algorithm works directly in the space of expected feature counts. In each iteration, it projects the expert feature expectations onto the convex set of feature expectations generated by previously evaluated policies to determine an orthogonal weight vector that serves as a candidate reward function. It then computes an optimal policy for that candidate reward via standard reinforcement learning, calculates the new policy feature expectations, and updates the projection. This cycle repeats until the distance between the expert feature expectations and the projected policy mixture drops below a specified threshold, guaranteeing that the resulting policy achieves performance close to that of the expert under the true reward.
1 item

