XGen 8K Model Beats Baselines At Long Text Tasks | | | | Turtles AI

XGen 8K Model Beats Baselines At Long Text Tasks
DukeRem8 July 2023
  XGen, a new large language model (LLM) from #Salesforce, has achieved state-of-the-art results on long-form tasks thanks to its ability to process up to 8,192 #tokens of text. However, the researchers caution that like other large models, #XGen faces potential risks of amplifying biases and generating inaccurate or harmful content. The researchers trained a series of 7 billion parameter models called XGen-7B with standard dense attention. They focused on maximizing the sequence length up to 8K tokens while training on up to 1.5 trillion tokens of data. On standard benchmarks, XGen achieves comparable or better results than other open-source language models of similar size. However, where XGen truly shines is on long-form tasks that require understanding and generating long sequences of text. The researchers found XGen performed better at summarizing long dialogues and screenplays over 6,000 tokens. It also achieved higher coherence and relevance when answering questions about Wikipedia texts over 4,000 tokens long. Its code generation capabilities were also stronger on tasks with long instructions. Despite these promising results, the researchers caution that like all large language models, XGen still faces risks of amplifying social biases, generating toxic or inaccurate content, and lacking diversity in its training data. The researchers hope their open-sourced codebase will help mitigate these issues and ensure AI benefits society as a whole. Have a look at our guides on LLMs, to better understand the basics.