Artificial Intelligence Value Alignment via Inverse Reinforcement Learning
Duim, João Lucas
O documento é disponibilizado pela fonte de origem, que mantém a versão integral e as condições de uso.
Resumo
This dissertation presents a comprehensive exploration of the field of artificial intelligence (AI) alignment, emphasizing the integration of human values and ethics into AI systems. The research synthesizes a broad range of academic literature to elucidate the principles, methodologies, challenges, and ethical implications inherent in AI alignment. Central to this discussion are the principles of robustness, interpretability, controllability, and ethicality, which are critical for the development of AI systems that are not only technically proficient but also ethically aligned with human values. The dissertation delves into the dual aspects of AI alignment: forward alignment, which focuses on embedding human values during the AI training phase, and backward alignment, emphasizing ongoing governance and verification post-deployment. A key challenge identified is the integration of complex and often subjective human values into computational models, highlighting limitations in current methodologies like Inverse Reinforcement Learning (IRL). The ethical and safety considerations in IRL are critically examined, underscoring the need for a balance between technological advancement and ethical integrity. The dissertation advocates for methodological advancements, hybrid approaches combining empirical data and ethical reasoning, addressing data biases, and establishing robust governance frameworks. Future research directions identified include methodological innovations, addressing data biases, and the need for interdisciplinary collaborations to tackle the multifaceted challenges of AI alignment. This research concludes that AI alignment is vital for addressing existential risks posed by AI, as it ensures AI development with human values and ethics, a key step in preventing AI technologies from diverging in potentially harmful ways. By emphasizing the principles of ethical alignment in AI systems, it contributes to mitigating the risks of unaligned powerful AIs and ensuring a safe, harmonious coexistence between AI and humanity.
Ficha do documento
- Tipo
- Outro
- Ano
- 2023
- Instituição
- Fundação Getulio Vargas
- Fonte
- Repositório da FGV
- Idioma
- Inglês
- Acesso
- Acesso aberto
- Identificador
- oai:repositorio.fgv.br:10438/35386
- Temas
- TecnologiaGovernança
Conteúdos relacionados
- DissertaçãoRobust unlearning via randomly initialized distillationFundação Getulio Vargas · 2026
- TeseExploring urban safety perception with vision-language models and semantically grounded image counterfactualsFundação Getulio Vargas · 2026
- DissertaçãoEmerging AI governance practices, challenges, and opportunities in industryFundação Getulio Vargas · 2024
- OutroRegulação de opacidade algorítmicaFundação Getulio Vargas · 2023