Logo
Outro

Artificial Intelligence Value Alignment via Inverse Reinforcement Learning

Duim, João Lucas

O documento é disponibilizado pela fonte de origem, que mantém a versão integral e as condições de uso.

Resumo

This dissertation presents a comprehensive exploration of the field of artificial intelligence (AI) alignment, emphasizing the integration of human values and ethics into AI systems. The research synthesizes a broad range of academic literature to elucidate the principles, methodologies, challenges, and ethical implications inherent in AI alignment. Central to this discussion are the principles of robustness, interpretability, controllability, and ethicality, which are critical for the development of AI systems that are not only technically proficient but also ethically aligned with human values. The dissertation delves into the dual aspects of AI alignment: forward alignment, which focuses on embedding human values during the AI training phase, and backward alignment, emphasizing ongoing governance and verification post-deployment. A key challenge identified is the integration of complex and often subjective human values into computational models, highlighting limitations in current methodologies like Inverse Reinforcement Learning (IRL). The ethical and safety considerations in IRL are critically examined, underscoring the need for a balance between technological advancement and ethical integrity. The dissertation advocates for methodological advancements, hybrid approaches combining empirical data and ethical reasoning, addressing data biases, and establishing robust governance frameworks. Future research directions identified include methodological innovations, addressing data biases, and the need for interdisciplinary collaborations to tackle the multifaceted challenges of AI alignment. This research concludes that AI alignment is vital for addressing existential risks posed by AI, as it ensures AI development with human values and ethics, a key step in preventing AI technologies from diverging in potentially harmful ways. By emphasizing the principles of ethical alignment in AI systems, it contributes to mitigating the risks of unaligned powerful AIs and ensuring a safe, harmonious coexistence between AI and humanity.

Ficha do documento

Tipo
Outro
Ano
2023
Instituição
Fundação Getulio Vargas
Idioma
Inglês
Acesso
Acesso aberto
Identificador
oai:repositorio.fgv.br:10438/35386

Conteúdos relacionados

Voltar à Biblioteca
Logo