LCFO: Long Context and Long Form Output Dataset and Benchmarking Article Swipe
YOU?
·
· 2024
· Open Access
·
· DOI: https://doi.org/10.48550/arxiv.2412.08268
This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across diverse domains. LCFO consists of long input documents (5k words average length), each of which comes with three summaries of different lengths (20%, 10%, and 5% of the input text), as well as approximately 15 questions and answers (QA) related to the input content. Notably, LCFO also provides alignments between specific QA pairs and corresponding summaries in 7 domains. The primary motivation behind providing summaries of different lengths is to establish a controllable framework for generating long texts from shorter inputs, i.e. summary expansion. To establish an evaluation metric framework for summarization and summary expansion, we provide human evaluation scores for human-generated outputs, as well as results from various state-of-the-art large language models (LLMs). GPT-4o-mini achieves best human scores among automatic systems in both summarization and summary expansion tasks (~ +10% and +20%, respectively). It even surpasses human output quality in the case of short summaries (~ +7%). Overall automatic metrics achieve low correlations with human evaluation scores (~ 0.4) but moderate correlation on specific evaluation aspects such as fluency and attribution (~ 0.6).
Related Topics
- Type
- preprint
- Language
- en
- Landing Page
- http://arxiv.org/abs/2412.08268
- https://arxiv.org/pdf/2412.08268
- OA Status
- green
- Related Works
- 10
- OpenAlex ID
- https://openalex.org/W4405301274
Raw OpenAlex JSON
- OpenAlex ID
-
https://openalex.org/W4405301274Canonical identifier for this work in OpenAlex
- DOI
-
https://doi.org/10.48550/arxiv.2412.08268Digital Object Identifier
- Title
-
LCFO: Long Context and Long Form Output Dataset and BenchmarkingWork title
- Type
-
preprintOpenAlex work type
- Language
-
enPrimary language
- Publication year
-
2024Year of publication
- Publication date
-
2024-12-11Full publication date if available
- Authors
-
Marta R. Costa‐jussà, Pierre Andrews, Mariano Coria Meglioli, Joy Chen, Jen-Hsiang Chuang, David C. Dale, Christophe Ropers, Alexandre Mourachko, Eduardo Sánchez, Holger Schwenk, Tuan Tran, Arina Turkatenko, Carleigh WoodList of authors in order
- Landing page
-
https://arxiv.org/abs/2412.08268Publisher landing page
- PDF URL
-
https://arxiv.org/pdf/2412.08268Direct link to full text PDF
- Open access
-
YesWhether a free full text is available
- OA status
-
greenOpen access status per OpenAlex
- OA URL
-
https://arxiv.org/pdf/2412.08268Direct OA link when available
- Concepts
-
Benchmarking, Context (archaeology), Computer science, Data science, Business, Geography, Archaeology, MarketingTop concepts (fields/topics) attached by OpenAlex
- Cited by
-
0Total citation count in OpenAlex
- Related works (count)
-
10Other works algorithmically related by OpenAlex
Full payload
| id | https://openalex.org/W4405301274 |
|---|---|
| doi | https://doi.org/10.48550/arxiv.2412.08268 |
| ids.doi | https://doi.org/10.48550/arxiv.2412.08268 |
| ids.openalex | https://openalex.org/W4405301274 |
| fwci | |
| type | preprint |
| title | LCFO: Long Context and Long Form Output Dataset and Benchmarking |
| biblio.issue | |
| biblio.volume | |
| biblio.last_page | |
| biblio.first_page | |
| topics[0].id | https://openalex.org/T11801 |
| topics[0].field.id | https://openalex.org/fields/22 |
| topics[0].field.display_name | Engineering |
| topics[0].score | 0.734499990940094 |
| topics[0].domain.id | https://openalex.org/domains/3 |
| topics[0].domain.display_name | Physical Sciences |
| topics[0].subfield.id | https://openalex.org/subfields/2212 |
| topics[0].subfield.display_name | Ocean Engineering |
| topics[0].display_name | Reservoir Engineering and Simulation Methods |
| is_xpac | False |
| apc_list | |
| apc_paid | |
| concepts[0].id | https://openalex.org/C86251818 |
| concepts[0].level | 2 |
| concepts[0].score | 0.9243104457855225 |
| concepts[0].wikidata | https://www.wikidata.org/wiki/Q816754 |
| concepts[0].display_name | Benchmarking |
| concepts[1].id | https://openalex.org/C2779343474 |
| concepts[1].level | 2 |
| concepts[1].score | 0.657362699508667 |
| concepts[1].wikidata | https://www.wikidata.org/wiki/Q3109175 |
| concepts[1].display_name | Context (archaeology) |
| concepts[2].id | https://openalex.org/C41008148 |
| concepts[2].level | 0 |
| concepts[2].score | 0.538227915763855 |
| concepts[2].wikidata | https://www.wikidata.org/wiki/Q21198 |
| concepts[2].display_name | Computer science |
| concepts[3].id | https://openalex.org/C2522767166 |
| concepts[3].level | 1 |
| concepts[3].score | 0.322051078081131 |
| concepts[3].wikidata | https://www.wikidata.org/wiki/Q2374463 |
| concepts[3].display_name | Data science |
| concepts[4].id | https://openalex.org/C144133560 |
| concepts[4].level | 0 |
| concepts[4].score | 0.22649729251861572 |
| concepts[4].wikidata | https://www.wikidata.org/wiki/Q4830453 |
| concepts[4].display_name | Business |
| concepts[5].id | https://openalex.org/C205649164 |
| concepts[5].level | 0 |
| concepts[5].score | 0.15446478128433228 |
| concepts[5].wikidata | https://www.wikidata.org/wiki/Q1071 |
| concepts[5].display_name | Geography |
| concepts[6].id | https://openalex.org/C166957645 |
| concepts[6].level | 1 |
| concepts[6].score | 0.0662308931350708 |
| concepts[6].wikidata | https://www.wikidata.org/wiki/Q23498 |
| concepts[6].display_name | Archaeology |
| concepts[7].id | https://openalex.org/C162853370 |
| concepts[7].level | 1 |
| concepts[7].score | 0.0660407543182373 |
| concepts[7].wikidata | https://www.wikidata.org/wiki/Q39809 |
| concepts[7].display_name | Marketing |
| keywords[0].id | https://openalex.org/keywords/benchmarking |
| keywords[0].score | 0.9243104457855225 |
| keywords[0].display_name | Benchmarking |
| keywords[1].id | https://openalex.org/keywords/context |
| keywords[1].score | 0.657362699508667 |
| keywords[1].display_name | Context (archaeology) |
| keywords[2].id | https://openalex.org/keywords/computer-science |
| keywords[2].score | 0.538227915763855 |
| keywords[2].display_name | Computer science |
| keywords[3].id | https://openalex.org/keywords/data-science |
| keywords[3].score | 0.322051078081131 |
| keywords[3].display_name | Data science |
| keywords[4].id | https://openalex.org/keywords/business |
| keywords[4].score | 0.22649729251861572 |
| keywords[4].display_name | Business |
| keywords[5].id | https://openalex.org/keywords/geography |
| keywords[5].score | 0.15446478128433228 |
| keywords[5].display_name | Geography |
| keywords[6].id | https://openalex.org/keywords/archaeology |
| keywords[6].score | 0.0662308931350708 |
| keywords[6].display_name | Archaeology |
| keywords[7].id | https://openalex.org/keywords/marketing |
| keywords[7].score | 0.0660407543182373 |
| keywords[7].display_name | Marketing |
| language | en |
| locations[0].id | pmh:oai:arXiv.org:2412.08268 |
| locations[0].is_oa | True |
| locations[0].source.id | https://openalex.org/S4306400194 |
| locations[0].source.issn | |
| locations[0].source.type | repository |
| locations[0].source.is_oa | True |
| locations[0].source.issn_l | |
| locations[0].source.is_core | False |
| locations[0].source.is_in_doaj | False |
| locations[0].source.display_name | arXiv (Cornell University) |
| locations[0].source.host_organization | https://openalex.org/I205783295 |
| locations[0].source.host_organization_name | Cornell University |
| locations[0].source.host_organization_lineage | https://openalex.org/I205783295 |
| locations[0].license | |
| locations[0].pdf_url | https://arxiv.org/pdf/2412.08268 |
| locations[0].version | submittedVersion |
| locations[0].raw_type | text |
| locations[0].license_id | |
| locations[0].is_accepted | False |
| locations[0].is_published | False |
| locations[0].raw_source_name | |
| locations[0].landing_page_url | http://arxiv.org/abs/2412.08268 |
| locations[1].id | doi:10.48550/arxiv.2412.08268 |
| locations[1].is_oa | True |
| locations[1].source.id | https://openalex.org/S4306400194 |
| locations[1].source.issn | |
| locations[1].source.type | repository |
| locations[1].source.is_oa | True |
| locations[1].source.issn_l | |
| locations[1].source.is_core | False |
| locations[1].source.is_in_doaj | False |
| locations[1].source.display_name | arXiv (Cornell University) |
| locations[1].source.host_organization | https://openalex.org/I205783295 |
| locations[1].source.host_organization_name | Cornell University |
| locations[1].source.host_organization_lineage | https://openalex.org/I205783295 |
| locations[1].license | |
| locations[1].pdf_url | |
| locations[1].version | |
| locations[1].raw_type | article |
| locations[1].license_id | |
| locations[1].is_accepted | False |
| locations[1].is_published | |
| locations[1].raw_source_name | |
| locations[1].landing_page_url | https://doi.org/10.48550/arxiv.2412.08268 |
| indexed_in | arxiv, datacite |
| authorships[0].author.id | https://openalex.org/A5074210163 |
| authorships[0].author.orcid | https://orcid.org/0000-0002-5703-520X |
| authorships[0].author.display_name | Marta R. Costa‐jussà |
| authorships[0].author_position | first |
| authorships[0].raw_author_name | Costa-jussà, Marta R. |
| authorships[0].is_corresponding | False |
| authorships[1].author.id | https://openalex.org/A5091619219 |
| authorships[1].author.orcid | https://orcid.org/0000-0001-6780-7798 |
| authorships[1].author.display_name | Pierre Andrews |
| authorships[1].author_position | middle |
| authorships[1].raw_author_name | Andrews, Pierre |
| authorships[1].is_corresponding | False |
| authorships[2].author.id | https://openalex.org/A5093470965 |
| authorships[2].author.orcid | |
| authorships[2].author.display_name | Mariano Coria Meglioli |
| authorships[2].author_position | middle |
| authorships[2].raw_author_name | Meglioli, Mariano Coria |
| authorships[2].is_corresponding | False |
| authorships[3].author.id | https://openalex.org/A5045909442 |
| authorships[3].author.orcid | |
| authorships[3].author.display_name | Joy Chen |
| authorships[3].author_position | middle |
| authorships[3].raw_author_name | Chen, Joy |
| authorships[3].is_corresponding | False |
| authorships[4].author.id | https://openalex.org/A5076500156 |
| authorships[4].author.orcid | https://orcid.org/0000-0002-7114-3418 |
| authorships[4].author.display_name | Jen-Hsiang Chuang |
| authorships[4].author_position | middle |
| authorships[4].raw_author_name | Chuang, Joe |
| authorships[4].is_corresponding | False |
| authorships[5].author.id | https://openalex.org/A5015520404 |
| authorships[5].author.orcid | https://orcid.org/0000-0002-7146-1440 |
| authorships[5].author.display_name | David C. Dale |
| authorships[5].author_position | middle |
| authorships[5].raw_author_name | Dale, David |
| authorships[5].is_corresponding | False |
| authorships[6].author.id | https://openalex.org/A5030185276 |
| authorships[6].author.orcid | |
| authorships[6].author.display_name | Christophe Ropers |
| authorships[6].author_position | middle |
| authorships[6].raw_author_name | Ropers, Christophe |
| authorships[6].is_corresponding | False |
| authorships[7].author.id | https://openalex.org/A5010669494 |
| authorships[7].author.orcid | |
| authorships[7].author.display_name | Alexandre Mourachko |
| authorships[7].author_position | middle |
| authorships[7].raw_author_name | Mourachko, Alexandre |
| authorships[7].is_corresponding | False |
| authorships[8].author.id | https://openalex.org/A5031287694 |
| authorships[8].author.orcid | https://orcid.org/0000-0002-8044-938X |
| authorships[8].author.display_name | Eduardo Sánchez |
| authorships[8].author_position | middle |
| authorships[8].raw_author_name | Sánchez, Eduardo |
| authorships[8].is_corresponding | False |
| authorships[9].author.id | https://openalex.org/A5015857371 |
| authorships[9].author.orcid | |
| authorships[9].author.display_name | Holger Schwenk |
| authorships[9].author_position | middle |
| authorships[9].raw_author_name | Schwenk, Holger |
| authorships[9].is_corresponding | False |
| authorships[10].author.id | https://openalex.org/A5035160416 |
| authorships[10].author.orcid | https://orcid.org/0000-0002-5132-6495 |
| authorships[10].author.display_name | Tuan Tran |
| authorships[10].author_position | middle |
| authorships[10].raw_author_name | Tran, Tuan |
| authorships[10].is_corresponding | False |
| authorships[11].author.id | https://openalex.org/A5115405368 |
| authorships[11].author.orcid | |
| authorships[11].author.display_name | Arina Turkatenko |
| authorships[11].author_position | middle |
| authorships[11].raw_author_name | Turkatenko, Arina |
| authorships[11].is_corresponding | False |
| authorships[12].author.id | https://openalex.org/A5112921393 |
| authorships[12].author.orcid | |
| authorships[12].author.display_name | Carleigh Wood |
| authorships[12].author_position | last |
| authorships[12].raw_author_name | Wood, Carleigh |
| authorships[12].is_corresponding | False |
| has_content.pdf | False |
| has_content.grobid_xml | False |
| is_paratext | False |
| open_access.is_oa | True |
| open_access.oa_url | https://arxiv.org/pdf/2412.08268 |
| open_access.oa_status | green |
| open_access.any_repository_has_fulltext | False |
| created_date | 2025-10-10T00:00:00 |
| display_name | LCFO: Long Context and Long Form Output Dataset and Benchmarking |
| has_fulltext | False |
| is_retracted | False |
| updated_date | 2025-11-06T06:51:31.235846 |
| primary_topic.id | https://openalex.org/T11801 |
| primary_topic.field.id | https://openalex.org/fields/22 |
| primary_topic.field.display_name | Engineering |
| primary_topic.score | 0.734499990940094 |
| primary_topic.domain.id | https://openalex.org/domains/3 |
| primary_topic.domain.display_name | Physical Sciences |
| primary_topic.subfield.id | https://openalex.org/subfields/2212 |
| primary_topic.subfield.display_name | Ocean Engineering |
| primary_topic.display_name | Reservoir Engineering and Simulation Methods |
| related_works | https://openalex.org/W4391375266, https://openalex.org/W2899084033, https://openalex.org/W2748952813, https://openalex.org/W4238897586, https://openalex.org/W435179959, https://openalex.org/W2619091065, https://openalex.org/W2059640416, https://openalex.org/W1490753184, https://openalex.org/W2284465472, https://openalex.org/W2291782699 |
| cited_by_count | 0 |
| locations_count | 2 |
| best_oa_location.id | pmh:oai:arXiv.org:2412.08268 |
| best_oa_location.is_oa | True |
| best_oa_location.source.id | https://openalex.org/S4306400194 |
| best_oa_location.source.issn | |
| best_oa_location.source.type | repository |
| best_oa_location.source.is_oa | True |
| best_oa_location.source.issn_l | |
| best_oa_location.source.is_core | False |
| best_oa_location.source.is_in_doaj | False |
| best_oa_location.source.display_name | arXiv (Cornell University) |
| best_oa_location.source.host_organization | https://openalex.org/I205783295 |
| best_oa_location.source.host_organization_name | Cornell University |
| best_oa_location.source.host_organization_lineage | https://openalex.org/I205783295 |
| best_oa_location.license | |
| best_oa_location.pdf_url | https://arxiv.org/pdf/2412.08268 |
| best_oa_location.version | submittedVersion |
| best_oa_location.raw_type | text |
| best_oa_location.license_id | |
| best_oa_location.is_accepted | False |
| best_oa_location.is_published | False |
| best_oa_location.raw_source_name | |
| best_oa_location.landing_page_url | http://arxiv.org/abs/2412.08268 |
| primary_location.id | pmh:oai:arXiv.org:2412.08268 |
| primary_location.is_oa | True |
| primary_location.source.id | https://openalex.org/S4306400194 |
| primary_location.source.issn | |
| primary_location.source.type | repository |
| primary_location.source.is_oa | True |
| primary_location.source.issn_l | |
| primary_location.source.is_core | False |
| primary_location.source.is_in_doaj | False |
| primary_location.source.display_name | arXiv (Cornell University) |
| primary_location.source.host_organization | https://openalex.org/I205783295 |
| primary_location.source.host_organization_name | Cornell University |
| primary_location.source.host_organization_lineage | https://openalex.org/I205783295 |
| primary_location.license | |
| primary_location.pdf_url | https://arxiv.org/pdf/2412.08268 |
| primary_location.version | submittedVersion |
| primary_location.raw_type | text |
| primary_location.license_id | |
| primary_location.is_accepted | False |
| primary_location.is_published | False |
| primary_location.raw_source_name | |
| primary_location.landing_page_url | http://arxiv.org/abs/2412.08268 |
| publication_date | 2024-12-11 |
| publication_year | 2024 |
| referenced_works_count | 0 |
| abstract_inverted_index.7 | 81 |
| abstract_inverted_index.a | 11, 95 |
| abstract_inverted_index.(~ | 153, 170, 182, 196 |
| abstract_inverted_index.15 | 58 |
| abstract_inverted_index.5% | 49 |
| abstract_inverted_index.It | 158 |
| abstract_inverted_index.QA | 75 |
| abstract_inverted_index.To | 108 |
| abstract_inverted_index.an | 110 |
| abstract_inverted_index.as | 54, 56, 127, 129, 192 |
| abstract_inverted_index.in | 80, 146, 164 |
| abstract_inverted_index.is | 92 |
| abstract_inverted_index.of | 28, 37, 43, 50, 89, 167 |
| abstract_inverted_index.on | 187 |
| abstract_inverted_index.to | 64, 93 |
| abstract_inverted_index.we | 119 |
| abstract_inverted_index.(5k | 32 |
| abstract_inverted_index.The | 83 |
| abstract_inverted_index.and | 6, 19, 48, 60, 77, 116, 149, 155, 194 |
| abstract_inverted_index.but | 184 |
| abstract_inverted_index.for | 15, 98, 114, 124 |
| abstract_inverted_index.low | 176 |
| abstract_inverted_index.the | 3, 51, 65, 165 |
| abstract_inverted_index.(QA) | 62 |
| abstract_inverted_index.+10% | 154 |
| abstract_inverted_index.0.4) | 183 |
| abstract_inverted_index.10%, | 47 |
| abstract_inverted_index.Form | 7 |
| abstract_inverted_index.LCFO | 26, 69 |
| abstract_inverted_index.Long | 4 |
| abstract_inverted_index.This | 0 |
| abstract_inverted_index.also | 70 |
| abstract_inverted_index.best | 140 |
| abstract_inverted_index.both | 147 |
| abstract_inverted_index.case | 166 |
| abstract_inverted_index.each | 36 |
| abstract_inverted_index.even | 159 |
| abstract_inverted_index.from | 102, 131 |
| abstract_inverted_index.i.e. | 105 |
| abstract_inverted_index.long | 29, 100 |
| abstract_inverted_index.such | 191 |
| abstract_inverted_index.well | 55, 128 |
| abstract_inverted_index.with | 40, 178 |
| abstract_inverted_index.(20%, | 46 |
| abstract_inverted_index.+20%, | 156 |
| abstract_inverted_index.+7%). | 171 |
| abstract_inverted_index.0.6). | 197 |
| abstract_inverted_index.among | 143 |
| abstract_inverted_index.comes | 39 |
| abstract_inverted_index.human | 121, 141, 161, 179 |
| abstract_inverted_index.input | 30, 52, 66 |
| abstract_inverted_index.large | 134 |
| abstract_inverted_index.novel | 12 |
| abstract_inverted_index.pairs | 76 |
| abstract_inverted_index.paper | 1 |
| abstract_inverted_index.short | 168 |
| abstract_inverted_index.tasks | 152 |
| abstract_inverted_index.texts | 101 |
| abstract_inverted_index.three | 41 |
| abstract_inverted_index.which | 38 |
| abstract_inverted_index.words | 33 |
| abstract_inverted_index.(LCFO) | 9 |
| abstract_inverted_index.Output | 8 |
| abstract_inverted_index.across | 23 |
| abstract_inverted_index.behind | 86 |
| abstract_inverted_index.metric | 112 |
| abstract_inverted_index.models | 136 |
| abstract_inverted_index.output | 162 |
| abstract_inverted_index.scores | 123, 142, 181 |
| abstract_inverted_index.text), | 53 |
| abstract_inverted_index.(LLMs). | 137 |
| abstract_inverted_index.Context | 5 |
| abstract_inverted_index.Overall | 172 |
| abstract_inverted_index.achieve | 175 |
| abstract_inverted_index.answers | 61 |
| abstract_inverted_index.aspects | 190 |
| abstract_inverted_index.average | 34 |
| abstract_inverted_index.between | 73 |
| abstract_inverted_index.diverse | 24 |
| abstract_inverted_index.fluency | 193 |
| abstract_inverted_index.gradual | 17 |
| abstract_inverted_index.inputs, | 104 |
| abstract_inverted_index.lengths | 45, 91 |
| abstract_inverted_index.metrics | 174 |
| abstract_inverted_index.primary | 84 |
| abstract_inverted_index.provide | 120 |
| abstract_inverted_index.quality | 163 |
| abstract_inverted_index.related | 63 |
| abstract_inverted_index.results | 130 |
| abstract_inverted_index.shorter | 103 |
| abstract_inverted_index.summary | 20, 106, 117, 150 |
| abstract_inverted_index.systems | 145 |
| abstract_inverted_index.various | 132 |
| abstract_inverted_index.Notably, | 68 |
| abstract_inverted_index.achieves | 139 |
| abstract_inverted_index.consists | 27 |
| abstract_inverted_index.content. | 67 |
| abstract_inverted_index.domains. | 25, 82 |
| abstract_inverted_index.language | 135 |
| abstract_inverted_index.length), | 35 |
| abstract_inverted_index.moderate | 185 |
| abstract_inverted_index.outputs, | 126 |
| abstract_inverted_index.presents | 2 |
| abstract_inverted_index.provides | 71 |
| abstract_inverted_index.specific | 74, 188 |
| abstract_inverted_index.assessing | 16 |
| abstract_inverted_index.automatic | 144, 173 |
| abstract_inverted_index.different | 44, 90 |
| abstract_inverted_index.documents | 31 |
| abstract_inverted_index.establish | 94, 109 |
| abstract_inverted_index.expansion | 21, 151 |
| abstract_inverted_index.framework | 14, 97, 113 |
| abstract_inverted_index.providing | 87 |
| abstract_inverted_index.questions | 59 |
| abstract_inverted_index.summaries | 42, 79, 88, 169 |
| abstract_inverted_index.surpasses | 160 |
| abstract_inverted_index.alignments | 72 |
| abstract_inverted_index.benchmark, | 10 |
| abstract_inverted_index.evaluation | 13, 111, 122, 180, 189 |
| abstract_inverted_index.expansion, | 118 |
| abstract_inverted_index.expansion. | 107 |
| abstract_inverted_index.generating | 99 |
| abstract_inverted_index.motivation | 85 |
| abstract_inverted_index.GPT-4o-mini | 138 |
| abstract_inverted_index.attribution | 195 |
| abstract_inverted_index.correlation | 186 |
| abstract_inverted_index.capabilities | 22 |
| abstract_inverted_index.controllable | 96 |
| abstract_inverted_index.correlations | 177 |
| abstract_inverted_index.approximately | 57 |
| abstract_inverted_index.corresponding | 78 |
| abstract_inverted_index.summarization | 18, 115, 148 |
| abstract_inverted_index.respectively). | 157 |
| abstract_inverted_index.human-generated | 125 |
| abstract_inverted_index.state-of-the-art | 133 |
| cited_by_percentile_year | |
| countries_distinct_count | 0 |
| institutions_distinct_count | 13 |
| citation_normalized_percentile |