Learning global optimization by deep reinforcement learning

Imagem de Miniatura

Data

2026-06-29

Lattes da Orientação Docente

Título da Revista

ISSN da Revista

Título de Volume

Editor

Resumo

Resumo em outro idioma

Learning to Optimize (L2O) is a growing field that employs a variety of machine learning (ML) methods to learn optimization algorithms automatically from data instead of developing handengineered algorithms that usually require hyperparameter tuning and problem-specific design. However, there are some barriers to adopting those learned optimizers in practice. For instance, they exhibit instability during training, poor generalization to problems outside the distribution, and lack scalability. Current research efforts suggest either improving L2O models or improving training techniques to overcome such hardships. We focus on the latter and propose to train a Deep Reinforcement Learning (Deep RL) agent to learn an optimization algorithm from training in a diverse set of benchmark functions. To this end, we propose a general framework for learning to optimize by reinforcement learning, which adapts training strategies used in other L2O approaches, such as curriculum learning and input normalization. We investigate the importance of these strategies through an ablation study and show that even though Deep RL, to the best of our knowledge, is not a well-explored theme in L2O, it provides a direct framework to learn an optimizer able to deal with the exploration-exploitation dilemma and that the applied techniques improved stability and generalization.

Descrição

Referência

SILVA FILHO, Moesio Wenceslau da. Learning global optimization by deep reinforcement learning. 2026. 15 f. Trabalho de Conclusão de Curso (Bacharelado em Ciência da Computação) – Departamento de Computação, Universidade Federal Rural de Pernambuco, Recife, 2026.

Identificador dARK

Avaliação

Revisão

Suplementado Por

Referenciado Por

Licença Creative Commons

Exceto quando indicado de outra forma, a licença deste item é descrita como openAccess