Shuo Ni1,3⋆, Tong Wang1⋆, Jing Zhang2,3†, Di Wang2,3, He Chen1, Haonan Guo3,4, Ning Zhang1,5†, Bo Du2,3†.
1 Beijing Institute of Technology, 2 Wuhan University, 3 Zhongguancun Academy, 4 State Key Laboratory of LIESMARS, Wuhan University, 5 Hong Kong Polytechnic University.
⋆ Equal contribution, † Corresponding author
Update | Abstract | Benchmark | Method | Usage | Citation | Statement
- [2026.09] 📄 Preprint under review. Source code released.
Vision-Language Models (VLMs) face a severe scale mismatch between broad scene context and micro-scale targets when analyzing ultra-high-resolution (UHR) Earth observation imagery — a phenomenon we term the resolution illusion. To benchmark this challenge, we introduce UHR-Micro, a benchmark of 11,072 instructions grounded in 1,212 UHR images spanning 16 tasks across grounding, counting, spatial reasoning, and fine-grained understanding. We further propose Micro-evidence Active Perception (MAP), a training-free method that constructs a spatially grounded evidence state from localized observations and supplements it when the current evidence is insufficient. Across two backbones, MAP improves UHR-Micro by 8.65 percentage points on average.
UHR-Micro evaluates whether VLMs can locate, interpret, and reason over task-relevant micro-evidence in UHR remote-sensing imagery. Each instruction is centered on an objectively verifiable target occupying less than 0.01% of the image area, spanning 16 tasks across four families.
MAP acquires localized native-resolution observations, records each observation with its source region, and composes a spatially grounded evidence state with the global context. When the evidence is insufficient, MAP selectively acquires additional observations to complete the state.
Three independent Python environments are required; see
ENVIRONMENT.md for the exact verified package versions.
Direct inference (Qwen3-VL-8B):
python -m direct.cli run \
--input-manifest /path/to/Test.jsonl \
--run-config /path/to/run_config.json \
--adapter-config direct/configs/qwen3_vl_8b.json \
--model-path $MAP_MODEL_PATH \
--image-root $MAP_DATASET/Datasource \
--gpus 0,1,2,3MAP (Qwen3-VL-8B, Validation):
MAP_RUN_DIR=/path/to/run \
MAP_DATASET=/path/to/UHR-Micro \
MAP_MANIFEST=/path/to/Validation.jsonl \
MAP_MODEL_PATH=/path/to/Qwen3-VL-8B-Instruct \
MAP_PHASE0=/path/to/phase0_contract.json \
MAP_EVALUATE=$(pwd)/evaluation/evaluate_formal.py \
python map/qwen3vl/run_validation.pyTest uses map/qwen3vl/run_test.py; InternVL3.5-8B uses
map/internvl/run_validation.py / run_test.py. All paths are supplied via
environment variables — see ENVIRONMENT.md.
If you find UHR-Micro helpful, please give a ⭐ and cite:
@misc{ni2026uhrmicro,
title={UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs},
author={Shuo Ni and Tong Wang and Jing Zhang and Di Wang and He Chen and Haonan Guo and Ning Zhang and Bo Du},
year={2026},
eprint={XXXX.XXXXX},
archivePrefix={arXiv},
primaryClass={cs.CV},
note={Under review}
}
For any other questions please contact Shuo Ni at shuoni@bit.edu.cn.
This project uses SAM2, Qwen3-VL, and InternVL. The benchmark is built on FAIR1Mv2, DOTAv2, SODA-A, and xView. Thanks for their wonderful work!


