Search CORE

1 research outputs found

StarCoder: may the source be with you!

Author: Abulkhanov Dmitry
Akiki Christopher
Allal Loubna Ben
Anderson Carolyn Jane
Bahdanau Dzmitry
Bhattacharyya Urvashi
Chim Jenny
Contractor Danish
Dao Tri
Davaadorj Mishig
de Vries Harm
Dehaene Olivier
Dey Manan
Ding Jennifer
Dolan-Gavitt Brendan
Ebert Jan
Fahmy Nour
Ferrandis Carlos Muñoz
Fried Daniel
Gontier Nicolas
Gu Alex
Guha Arjun
Hughes Sean
Jernite Yacine
Kocetkov Denis
Kunakov Maxim
Lamy-Poirier Joel
Lee Tony
Li Jia
Li Raymond
Lipkin Benjamin
Liu Qian
Luccioni Sasha
Marone Marc
Meade Nicholas
Mishra Mayank
Monteiro João
Mou Chenghao
Muennighoff Niklas
Murthy Rudra
Oblokulov Muhtasham
Patel Siva Sankalp
Reddy Siva
Robinson Jennifer
Romero Manuel
Schlesinger Claire
Schoelkopf Hailey
Shliazhko Oleh
Singh Swayam
Stillerman Jason
Timor Nadav
Umapathi Logesh Kumar
Villegas Paulo
von Werra Leandro
Wang Thomas
Wang Zhiruo
Wolf Thomas
Yee Ming-Ho
Yu Wenhao
Zebaze Armel
Zhang Zhihan
Zhdanov Fedor
Zheltonozhskii Evgenii
Zhu Jian
Zhuo Terry Yue
Zi Yangtian
Zocca Marco
Publication venue
Publication date: 09/05/2023
Field of study

The BigCode community, an open-scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder and StarCoderBase: 15.5B parameter models with 8K context length, infilling capabilities and fast large-batch inference enabled by multi-query attention. StarCoderBase is trained on 1 trillion tokens sourced from The Stack, a large collection of permissively licensed GitHub repositories with inspection tools and an opt-out process. We fine-tuned StarCoderBase on 35B Python tokens, resulting in the creation of StarCoder. We perform the most comprehensive evaluation of Code LLMs to date and show that StarCoderBase outperforms every open Code LLM that supports multiple programming languages and matches or outperforms the OpenAI code-cushman-001 model. Furthermore, StarCoder outperforms every model that is fine-tuned on Python, can be prompted to achieve 40\% pass@1 on HumanEval, and still retains its performance on other programming languages. We take several important steps towards a safe open-access model release, including an improved PII redaction pipeline and a novel attribution tracing tool, and make the StarCoder models publicly available under a more commercially viable version of the Open Responsible AI Model license

arXiv.org e-Print Archive