An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination
  Evaluation

Gu, Yukai; Jia, Haitao; Sang, Jitao; Wang, Junyang; Wang, Yuhang; Xu, Guohai; Yan, Ming; Zhang, Ji; Zhang, Jing

An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Authors: Yukai Gu
Haitao Jia
Jitao Sang
Junyang Wang
Yuhang Wang
Guohai Xu
Ming Yan
Ji Zhang
Jing Zhang
Publication date: 13 November 2023
Publisher

Abstract

Despite making significant progress in multi-modal tasks, current Multi-modal Large Language Models (MLLMs) encounter the significant challenge of hallucination, which may lead to harmful consequences. Therefore, evaluating MLLMs' hallucinations is becoming increasingly important in model improvement and practical application deployment. Previous works are limited in high evaluation costs (e.g., relying on humans or advanced LLMs) and insufficient evaluation dimensions (e.g., types of hallucination and task). In this paper, we propose an LLM-free multi-dimensional benchmark AMBER, which can be used to evaluate both generative task and discriminative task including object existence, object attribute and object relation hallucination. Based on AMBER, we design a low-cost and efficient evaluation pipeline. Additionally, we conduct a comprehensive evaluation and detailed analysis of mainstream MLLMs including GPT-4V(ision), and also give guideline suggestions for mitigating hallucinations. The data and code of AMBER are available at https://github.com/junyangwang0410/AMBER.Comment: 11 pages, 4 figure

Similar works

Full text

Available Versions

arXiv.org e-Print Archive

oai:arXiv.org:2311.07397

Last time updated on 10/02/2024