dc.contributor.author | Jung, T. | |
dc.contributor.author | Polani, D. | |
dc.date.accessioned | 2011-10-20T11:01:11Z | |
dc.date.available | 2011-10-20T11:01:11Z | |
dc.date.issued | 2007 | |
dc.identifier.citation | Jung , T & Polani , D 2007 , Kernelizing LSPE λ . in Procs of the 2007 Symposium on Approximate Dynamic Programming & Reinforcement Learning (ADPRL 2007) . vol. 2007 , Institute of Electrical and Electronics Engineers (IEEE) , pp. 338-345 . | |
dc.identifier.isbn | 1-4244-0706-0 | |
dc.identifier.other | PURE: 425284 | |
dc.identifier.other | PURE UUID: 3c26bf2c-6982-44b3-95fa-72d79185bbf2 | |
dc.identifier.other | dspace: 2299/1920 | |
dc.identifier.other | Scopus: 34548765672 | |
dc.identifier.other | ORCID: /0000-0002-3233-5847/work/86098100 | |
dc.identifier.uri | http://hdl.handle.net/2299/6735 | |
dc.description.abstract | We propose the use of kernel-based methods as underlying function approximator in the least-squares based policy evaluation framework of LSPE(λ) and LSTD(λ). In particular we present the ‘kernelization’ of model-free LSPE(λ). The ‘kernelization’ is computationally made possible by using the subset of regressors approximation, which approximates the kernel using a vastly reduced number of basis functions. The core of our proposed solution is an efficient recursive implementation with automatic supervised selection of the relevant basis functions. The LSPE method is well-suited for optimistic policy iteration and can thus be used in the context of online reinforcement learning. We use the high-dimensional Octopus benchmark to demonstrate this. | en |
dc.language.iso | eng | |
dc.publisher | Institute of Electrical and Electronics Engineers (IEEE) | |
dc.relation.ispartof | Procs of the 2007 Symposium on Approximate Dynamic Programming & Reinforcement Learning (ADPRL 2007) | |
dc.title | Kernelizing LSPE λ | en |
dc.contributor.institution | Centre for Computer Science and Informatics Research | |
dc.contributor.institution | Adaptive Systems | |
dc.contributor.institution | Department of Computer Science | |
dc.contributor.institution | School of Physics, Engineering & Computer Science | |
dc.contributor.institution | Centre for Future Societies Research | |
rioxxterms.version | VoR | |
rioxxterms.type | Other | |
herts.preservation.rarelyaccessed | true | |