Compressing Prompts in LLMs | | | | Turtles AI
Compressing Prompts in LLMs
DukeRem23 April 2023
In a new paper, researchers Jesse Mu, Xiang Lisa Li, Noah Goodman (#Stanford #University) have proposed a framework for #prompt #compression in large language models (#LLM) called "#gisting". While prompts are an efficient way to utilize the #multitask capabilities of LMs, they occupy valuable space in the input context window, and re-encoding the same prompt is computationally inefficient. #Finetuning and distillation methods can specialize LMs without prompting, but require retraining for each task. Gisting, on the other hand, trains an LM to compress prompts into smaller sets of "gist" tokens which can be reused for compute efficiency. This method enables up to 26x compression of prompts, resulting in up to 40% FLOPs reductions, 4.2% wall time speedups, storage savings, and minimal loss in output quality. Gisting can be seen as a modified form of instruction finetuning or a method for (meta-)context distillation of an LM. The researchers believe that gisting opens up several promising directions for follow-up work, including parameter-efficient gisting and exploring longer prompt compression.
