Skip to content

Run evalutation of model on Dharmabench QUDT benchmark #413

Description

@karmatai

Description

The current finetuned scores really low on quotation detection. As I found there already exist a quotation benchmark developed by Dharmabench, it would be better to see how good our models are doing before moving into further implementations.

Subtask

  • Load the current model checkpoint.
  • Load the Dharmabench QUDT benchmark.
  • Run evaluation script
  • Document the result

Reviewer

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

High

Projects

Status
In Progress

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions