Numerical precision is critical in financial NLP, yet embedding-based semantic similarity metrics exhibit numerical blindness—failing to distinguish contradictory values within similar contexts. We introduce NASH (Numerically Aware Scoring Heuristic), a model-agnostic metric that decouples numerical verification from textual semantic evaluation through a three-stage pipeline: (1) modal separation via numeric masking, (2) dual-channel similarity estimation through masked-text similarity and context-aware numeric alignment, and (3) IDF-weighted aggregation. NASH functions as a drop-in enhancement to existing embedding-based metrics. Validated on our proposed NumFinE financial numerical evaluation benchmark and established semantic similarity datasets (STS-B, Financial-STS), NASH achieves substantial improvements in numerical sensitivity (up to +159.6% on listwise ranking) while preserving general semantic performance, establishing a reliable standard for numeracy-aware evaluation.
This repository provides a paper-aligned public implementation of NASH along with the evaluation datasets used in our experiments.
nash_public_release/
├── nash/
├── scripts/
├── data/
├── run_triplet.sh
├── run_crosspair.sh
├── run_listwise.sh
├── requirements.txt
└── README.md
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtWe provide executable scripts for all evaluation settings:
bash run_triplet.sh
bash run_crosspair.sh
bash run_listwise.shPlease refer to these scripts for detailed configurations.
- This release follows the pipeline described in the paper.
- The dataset used in our experiments is included in the repository.
- The implementation is intended for reproducibility and evaluation purposes.