Search CORE

552 research outputs found

Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT

Author: Hu Mingzhe
Li Yuheng
Pan Shaoyan
Yang Xiaofeng
Publication venue
Publication date: 10/04/2023
Field of study

In this paper, we aimed to provide a review and tutorial for researchers in the field of medical imaging using language models to improve their tasks at hand. We began by providing an overview of the history and concepts of language models, with a special focus on large language models. We then reviewed the current literature on how language models are being used to improve medical imaging, emphasizing different applications such as image captioning, report generation, report classification, finding extraction, visual question answering, interpretable diagnosis, and more for various modalities and organs. The ChatGPT was specially highlighted for researchers to explore more potential applications. We covered the potential benefits of accurate and efficient language models for medical imaging analysis, including improving clinical workflow efficiency, reducing diagnostic errors, and assisting healthcare professionals in providing timely and accurate diagnoses. Overall, our goal was to bridge the gap between language models and medical imaging and inspire new ideas and innovations in this exciting area of research. We hope that this review paper will serve as a useful resource for researchers in this field and encourage further exploration of the possibilities of language models in medical imaging

arXiv.org e-Print Archive

Artificial General Intelligence for Medical Imaging

Author: Li Gang
Li Quanzheng
Li Xiang
Liu Jun
Liu Tianming
Liu Wei
Liu Zhengliang
Shen Dinggang
Wu Zihao
Yan Pingkuan
Yuan Yixuan
Zhang Lu
Zhao Lin
Zhu Dajiang
Publication venue
Publication date: 08/06/2023
Field of study

In this review, we explore the potential applications of Artificial General Intelligence (AGI) models in healthcare, focusing on foundational Large Language Models (LLMs), Large Vision Models, and Large Multimodal Models. We emphasize the importance of integrating clinical expertise, domain knowledge, and multimodal capabilities into AGI models. In addition, we lay out key roadmaps that guide the development and deployment of healthcare AGI models. Throughout the review, we provide critical perspectives on the potential challenges and pitfalls associated with deploying large-scale AGI models in the medical field. This comprehensive review aims to offer insights into the future implications of AGI in medical imaging, healthcare and beyond

arXiv.org e-Print Archive

Neural Natural Language Generation: A Survey on Multilinguality, Multimodality, Controllability and Learning

Author: Apostol Elena-Simona
Babii Andrii
Berend Gábor
Calixto Iacer
Erdem Aykut
Erdem Erkut
Frank Anette
Gatt Albert
Korvel Grăzina
Kuyu Menekse
Lloret Elena
Martinčić-Ipšić Sanda
Parcalabescu Letitia
Truică Ciprian-Octavian
Turuta Oleksii
Yagcioglu Semih
Šandrih Branislava
Publication venue: 'AI Access Foundation'
Publication date: 06/04/2022
Field of study

Developing artificial learning systems that can understand and generate natural language has been one of the long-standing goals of artificial intelligence. Recent decades have witnessed an impressive progress on both of these problems, giving rise to a new family of approaches. Especially, the advances in deep learning over the past couple of years have led to neural approaches to natural language generation (NLG). These methods combine generative language learning techniques with neural-networks based frameworks. With a wide range of applications in natural language processing, neural NLG (NNLG) is a new and fast growing field of research. In this state-of-the-art report, we investigate the recent developments and applications of NNLG in its full extent from a multidimensional view, covering critical perspectives such as multimodality, multilinguality, controllability and learning strategies. We summarize the fundamental building blocks of NNLG approaches from these aspects and provide detailed reviews of commonly used preprocessing steps and basic neural architectures. This report also focuses on the seminal applications of these NNLG models such as machine translation, description generation, automatic speech recognition, abstractive summarization, text simplification, question answering and generation, and dialogue generation. Finally, we conclude with a thorough discussion of the described frameworks by pointing out some open research directions.This work has been partially supported by the European Commission ICT COST Action “Multi-task, Multilingual, Multi-modal Language Generation” (CA18231). AE was supported by BAGEP 2021 Award of the Science Academy. EE was supported in part by TUBA GEBIP 2018 Award. BP is in in part funded by Independent Research Fund Denmark (DFF) grant 9063-00077B. IC has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant agreement No 838188. EL is partly funded by Generalitat Valenciana and the Spanish Government throught projects PROMETEU/2018/089 and RTI2018-094649-B-I00, respectively. SMI is partly funded by UNIRI project uniri-drustv-18-20. GB is partly supported by the Ministry of Innovation and the National Research, Development and Innovation Office within the framework of the Hungarian Artificial Intelligence National Laboratory Programme. COT is partially funded by the Romanian Ministry of European Investments and Projects through the Competitiveness Operational Program (POC) project “HOLOTRAIN” (grant no. 29/221 ap2/07.04.2020, SMIS code: 129077) and by the German Academic Exchange Service (DAAD) through the project “AWAKEN: content-Aware and netWork-Aware faKE News mitigation” (grant no. 91809005). ESA is partially funded by the German Academic Exchange Service (DAAD) through the project “Deep-Learning Anomaly Detection for Human and Automated Users Behavior” (grant no. 91809358)

PharmacyGPT: The AI Pharmacist

Author: Chen Xianyan
Dai Haixing
Hu Mengxuan
Li Sheng
Liu Tianming
Liu Zhengliang
Murray Brian
Shen Ye
Sikora Andrea
Wu Zihao
Zhang Tianyi
Zhao Bokai
Zhao Lin
Publication venue
Publication date: 20/07/2023
Field of study

In this study, we introduce PharmacyGPT, a novel framework to assess the capabilities of large language models (LLMs) such as ChatGPT and GPT-4 in emulating the role of clinical pharmacists. Our methodology encompasses the utilization of LLMs to generate comprehensible patient clusters, formulate medication plans, and forecast patient outcomes. We conduct our investigation using real data acquired from the intensive care unit (ICU) at the University of North Carolina Chapel Hill (UNC) Hospital. Our analysis offers valuable insights into the potential applications and limitations of LLMs in the field of clinical pharmacy, with implications for both patient care and the development of future AI-driven healthcare solutions. By evaluating the performance of PharmacyGPT, we aim to contribute to the ongoing discourse surrounding the integration of artificial intelligence in healthcare settings, ultimately promoting the responsible and efficacious use of such technologies

arXiv.org e-Print Archive

A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Author: Beirami Ahmad
Chen Shuo
Gu Jindong
Han Zhen
He Bailan
Liao Ruotong
Qin Yao
Torr Philip
Tresp Volker
Zhang Gengyuan
Publication venue
Publication date: 24/07/2023
Field of study

Prompt engineering is a technique that involves augmenting a large pre-trained model with task-specific hints, known as prompts, to adapt the model to new tasks. Prompts can be created manually as natural language instructions or generated automatically as either natural language instructions or vector representations. Prompt engineering enables the ability to perform predictions based solely on prompts without updating model parameters, and the easier application of large pre-trained models in real-world tasks. In past years, Prompt engineering has been well-studied in natural language processing. Recently, it has also been intensively studied in vision-language modeling. However, there is currently a lack of a systematic overview of prompt engineering on pre-trained vision-language models. This paper aims to provide a comprehensive survey of cutting-edge research in prompt engineering on three types of vision-language models: multimodal-to-text generation models (e.g. Flamingo), image-text matching models (e.g. CLIP), and text-to-image generation models (e.g. Stable Diffusion). For each type of model, a brief model summary, prompting methods, prompting-based applications, and the corresponding responsibility and integrity issues are summarized and discussed. Furthermore, the commonalities and differences between prompting on vision-language models, language models, and vision models are also discussed. The challenges, future directions, and research opportunities are summarized to foster future research on this topic

arXiv.org e-Print Archive

The Use of ChatGPT for Generating Scientific Citations : An experiment

Author: Bergman Jussi
Publication venue
Publication date: 21/08/2023
Field of study

This research examined the capabilities of ChatGPT in producing a numbered list of references for a range of topics and the accuracy of each reference through manual evaluation. Results suggest moderate levels of precision in generating reference lists, with 55% accuracy in titles, 43% in authors, 44% in sources, and 54% in overall relevance. Based on the relatively low accuracy of the generated references, this study introduced and applied a novel "Reverse Order Method". This method involves generating a list of references, manually validating each, and then instructing ChatGPT to compose a theoretical introduction based on the validated references alone. It's implied that the model's precise reproduction of a reference demonstrates repeated exposure and understanding of its content, enabling reliable citation in the final text. All final texts for all topics were evaluated as convincing and good quality scientific text with citations in place. Though the assessment of the final texts was purely subjective, the study suggest the promising utility of the Reverse Order Method in crafting scientific texts using ChatGPT 3.5. The study underscores the potential of AI tools like ChatGPT in scientific writing, emphasising the role of manual validation in improving precision and the careful use of AI-generated references. To enhance the understanding of the results, the study further explored the intricate inner workings of ChatGPT, concentrating on its 'transformer' architecture and the pursued 'learning objectives'. The study offered intuition by exploring the essential principles of language comprehension. A graphical representation, built upon existing research, was employed to illuminate this complex procedure

Large Language Models for Robotics: A Survey

Author: Gan Wensheng
Liu Ning
Wang Yongheng
Yu Philip S.
Zeng Fanlong
Publication venue
Publication date: 13/11/2023
Field of study

The human ability to learn, generalize, and control complex manipulation tasks through multi-modality feedback suggests a unique capability, which we refer to as dexterity intelligence. Understanding and assessing this intelligence is a complex task. Amidst the swift progress and extensive proliferation of large language models (LLMs), their applications in the field of robotics have garnered increasing attention. LLMs possess the ability to process and generate natural language, facilitating efficient interaction and collaboration with robots. Researchers and engineers in the field of robotics have recognized the immense potential of LLMs in enhancing robot intelligence, human-robot interaction, and autonomy. Therefore, this comprehensive review aims to summarize the applications of LLMs in robotics, delving into their impact and contributions to key areas such as robot control, perception, decision-making, and path planning. We first provide an overview of the background and development of LLMs for robotics, followed by a description of the benefits of LLMs for robotics and recent advancements in robotics models based on LLMs. We then delve into the various techniques used in the model, including those employed in perception, decision-making, control, and interaction. Finally, we explore the applications of LLMs in robotics and some potential challenges they may face in the near future. Embodied intelligence is the future of intelligent science, and LLMs-based robotics is one of the promising but challenging paths to achieve this.Comment: Preprint. 4 figures, 3 table

arXiv.org e-Print Archive

Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation

Author: Gatt Albert
Krahmer Emiel
Publication venue
Publication date: 01/01/2017
Field of study

This paper surveys the current state of the art in Natural Language Generation (NLG), defined as the task of generating text or speech from non-linguistic input. A survey of NLG is timely in view of the changes that the field has undergone over the past decade or so, especially in relation to new (usually data-driven) methods, as well as new applications of NLG technology. This survey therefore aims to (a) give an up-to-date synthesis of research on the core tasks in NLG and the architectures adopted in which such tasks are organised; (b) highlight a number of relatively recent research topics that have arisen partly as a result of growing synergies between NLG and other areas of artificial intelligence; (c) draw attention to the challenges in NLG evaluation, relating them to similar challenges faced in other areas of Natural Language Processing, with an emphasis on different evaluation methods and the relationships between them.Comment: Published in Journal of AI Research (JAIR), volume 61, pp 75-170. 118 pages, 8 figures, 1 tabl

arXiv.org e-Print Archive

Tilburg University Repository