Bootstrap your own latent: A new approach to self-supervised learning JB Grill, F Strub, F Altché, C Tallec, PH Richemond, E Buchatskaya, ... arXiv preprint arXiv:2006.07733, 2020 | 186 | 2020 |
Joint semantic utterance classification and slot filling with recursive neural networks D Guo, G Tur, W Yih, G Zweig 2014 IEEE Spoken Language Technology Workshop (SLT), 554-559, 2014 | 128 | 2014 |
Agent57: Outperforming the atari human benchmark AP Badia, B Piot, S Kapturowski, P Sprechmann, A Vitvitskyi, ZD Guo, ... International Conference on Machine Learning, 507-517, 2020 | 78 | 2020 |
Using options and covariance testing for long horizon off-policy policy evaluation ZD Guo, PS Thomas, E Brunskill arXiv preprint arXiv:1703.03453, 2017 | 30 | 2017 |
Neural predictive belief representations ZD Guo, MG Azar, B Piot, BA Pires, R Munos arXiv preprint arXiv:1811.06407, 2018 | 26 | 2018 |
Never give up: Learning directed exploration strategies AP Badia, P Sprechmann, A Vitvitskyi, D Guo, B Piot, S Kapturowski, ... arXiv preprint arXiv:2002.06038, 2020 | 24 | 2020 |
Concurrent pac rl Z Guo, E Brunskill Proceedings of the AAAI Conference on Artificial Intelligence 29 (1), 2015 | 19 | 2015 |
A pac rl algorithm for episodic pomdps ZD Guo, S Doroudi, E Brunskill Artificial Intelligence and Statistics, 510-518, 2016 | 15 | 2016 |
Bootstrap latent-predictive representations for multitask reinforcement learning ZD Guo, BA Pires, B Piot, JB Grill, F Altché, R Munos, MG Azar International Conference on Machine Learning, 3875-3886, 2020 | 14 | 2020 |
Pac continuous state online multitask reinforcement learning with identification Y Liu, Z Guo, E Brunskill Proceedings of the 2016 International Conference on Autonomous Agents …, 2016 | 7 | 2016 |
Never give up: Learning directed exploration strategies A Puigdomènech Badia, P Sprechmann, A Vitvitskyi, D Guo, B Piot, ... arXiv e-prints, arXiv: 2002.06038, 2020 | 6 | 2020 |
Sample efficient feature selection for factored mdps ZD Guo, E Brunskill arXiv preprint arXiv:1703.03454, 2017 | 6 | 2017 |
Agent57: Outperforming the Atari Human Benchmark A Puigdomènech Badia, B Piot, S Kapturowski, P Sprechmann, ... arXiv e-prints, arXiv: 2003.13350, 2020 | 3 | 2020 |
Sample efficient learning with feature selection for factored MDPs ZD Guo, E Brunskill Proceedings of the 14th European Workshop on Reinforcement Learning. EWRL, 2018 | 3 | 2018 |
Directed exploration for reinforcement learning ZD Guo, E Brunskill arXiv preprint arXiv:1906.07805, 2019 | 2 | 2019 |
Using options for long-horizon off-policy evaluation ZD Guo, PS Thomas, E Brunskill arXiv preprint arXiv:1703.03453, 2017 | 2 | 2017 |
Geometric Entropic Exploration ZD Guo, MG Azar, A Saade, S Thakoor, B Piot, BA Pires, M Valko, ... arXiv preprint arXiv:2101.02055, 2021 | 1 | 2021 |
Directed exploration for improved sample efficiency in reinforcement learning ZD Guo Google, 2019 | 1 | 2019 |