Description
The current finetuned scores really low on quotation detection. As I found there already exist a quotation benchmark developed by Dharmabench, it would be better to see how good our models are doing before moving into further implementations.
Subtask
Reviewer
Description
The current finetuned scores really low on quotation detection. As I found there already exist a quotation benchmark developed by Dharmabench, it would be better to see how good our models are doing before moving into further implementations.
Subtask
Reviewer