Sara Rosenthal, Pepa Atanasova, et al.
ACL-IJCNLP 2021
We present MTRAG-UN, a benchmark for exploring open challenges in multi-turn retrieval augment generation, a popular use of large language models. We release a benchmark of 666 tasks from 666 conversations containing over 2,800 conversation turns across 6 domains with accompanying corpora. Our experiments show that retrieval and generation models continue to struggle on conversations with UNanswerable, UNderspecified, and NONstandalone questions and UNclear responses.
Sara Rosenthal, Pepa Atanasova, et al.
ACL-IJCNLP 2021
Sola Shirai, Kavitha Srinivas, et al.
ACL 2026
Siya Kunde, Stephanie Houde, et al.
IUI 2025
Qiushi Huang, Xubo Liu, et al.
ACL 2024