SENSA: Unsupervised Rule Generation for More Robust Textual Summarization
Abstract
LLM summarization/compaction has become an increasingly important use case of LLMs as context sizes increase along with token prices. We introduce SENSA (Summarize, Evaluate, Nudge, Summarize Again), a novel summarization method that uses lightweight domain-aware preprocessing to steer summarization into attending more strongly to important information that is commonly lost during the summarization process. We demonstrate that in most domains, summarization errors can be described by a low-dimensional set of categories, and rules based on these categories can be used to guide summarization within that domain to be more resilient against information loss while keeping the same tight summarization budget. We demonstrate SENSA on both technical and non-technical domains, and show it beats SOTA compression on information retention for a wide range of domains and practical compression budgets. SENSA performs especially well on technical fields like coding, where relative improvement in compression information retention can be as high as 100% compared to the next best SOTA method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.