The Pile: An 800GB dataset of diverse text for language modeling (2020) by from on 2023-07-11 18:19 (#6CWSB) Comments