| layout |
home |
| title |
MARINA Dataset |
| description |
A large labeled dataset for underwater acoustic target recognition. |
| hero |
| background_image |
subheadline |
headline |
text |
buttons |
/assets/images/coming-soon-background.jpg |
Open benchmark · Underwater Acoustics |
<strong>UniqueShip:</strong> Large, public underwater acoustic target recognition (UATR) datasets for ships |
2,460 hours of ship-radiated noise from 4,218 unique vessels, split by vessel ID so no ship appears in both training and test split. Sourced from the Ocean Networks Canada (ONC) repository. |
| label |
url |
style |
Browse releases |
#releases |
primary |
|
| label |
url |
style |
Read the overview |
#overview |
outline |
|
|
|
| overview |
| subheadline |
headline |
paragraphs |
cta_section |
stats |
Overview |
Built to train generalizable UATR models using leakproof splits |
UniqueShip pairs hydrophone recordings from seven ONC deployments in the Strait of Georgia (May 2016 – November 2023) with AIS vessel tracking data. Each 5-second sample is labeled with its vessel class and 17 AIS metadata fields. Unlike earlier ONC-based datasets, every split keeps each vessel in a single partition and groups background audio by day, so test accuracy reflects performance on ships the model has never heard. |
|
| paragraphs |
buttons |
<strong>The paper</strong> describes the dataset in more detail and includes results with and without data leakage, ablations on vessel diversity vs. audio duration, a metadata analysis, and additional baseline results. |
|
| label |
url |
target |
style |
Read the paper |
|
_blank |
primary |
|
|
|
| value |
label |
3,437 |
Hours of ship & background audio |
|
| value |
label |
4,218 |
Unique vessels |
|
| value |
label |
11 |
Vessel classes |
|
| value |
label |
2.5M |
5-second recordings |
|
|
|
| difference |
| subheadline |
headline |
items |
What makes it different? |
Larger, more diverse, and free of data leakage that inflates other benchmarks |
| title |
text |
Leak-free splits |
Vessel audio is grouped by MMSI and background by day instead of random splitting. On previous datasets, random splitting inflated accuracy by 10–48 points. |
|
| title |
text |
Largest open ONC dataset |
The balanced benchmark subset alone has 4× the audio and 12× the vessels of DeepShip, and 70% more audio than the unbalanced Oceanship dataset. |
|
| title |
text |
Rich AIS metadata |
17 fields per sample, including MMSI, distance to hydrophone, speed, course, length, beam, draught, and navigation status. |
|
| title |
text |
Ready-made splits |
Choose anything from a 25-hour quick-start subset to the full 3,437-hour corpus, with five 80/10/10 folds. |
|
| title |
text |
Baselines included |
MobileNetV3, ViT-B/16, and SwinV2 with STFT and Mel inputs. The best result is 66.5% accuracy (Swin + Mel). |
|
| title |
text |
Cleaner background class |
8km ship-free radius ensures quieter ambient samples for the background class |
|
|
|
| releases |
| subheadline |
headline |
text |
cta_section |
items |
Data releases |
Current dataset splits |
All splits are vessel-disjoint and include per-sample AIS metadata. Samples are 5-second clips at 20 kHz; full-length recordings are available through the codebase. Each folder contains several zipped folders, which all must be unzipped. The labels and metadata for the datasets are given in the CSV files, where each row provides the relative path of the audio/spectrogram and its corresponding metadata/label. Please follow the leakproof folds for best standardization and benchmarking across multiple models. |
| paragraphs |
Data currently provided through Google Drive links, and please <strong>request Google Drive permission to access the current splits.</strong> The larger splits ("Main 5, Unbalanced") will be uploaded soon, with links to AWS. Please feel free to message if currently waiting. |
|
|
| title |
image |
image_alt |
description |
size |
views |
date |
buttons |
5 Class - Balanced |
/assets/images/5shiptypes.png |
Aerial view of five vessel classes tracked in open water |
Contains the main 5 classes (Tug/Tow, Tanker, Passengership, Cargo) and balances the total audio for each class such that they are equal. Current version = 1.0 |
92 GB (Unzipped), 64.8 GB (Zipped) |
213h, 3175 vessels, 5 classes |
Released September 2026 |
| label |
url |
target |
style |
Download audio |
|
_blank |
primary |
|
| label |
url |
target |
style |
Download spectrograms |
|
_blank |
primary |
|
|
|
| title |
image |
image_alt |
description |
size |
views |
date |
buttons |
12 Class - 5 Hours Each |
/assets/images/moreshiptypes.png |
Diverse vessels tracked in a busy coastal shipping channel |
Contains all ship classes and balances the total audio such that it is 5 hours each class. Current version = 1.0 |
26 GB (Unzipped), 17.8 GB (Zipped) |
60h, 4218 vessels, 12 classes |
Released September 2026 |
| label |
url |
target |
style |
Download audio |
|
_blank |
primary |
|
| label |
url |
target |
style |
Download spectrograms |
|
_blank |
primary |
|
|
|
|
|
| leaderboards |
| subheadline |
headline |
status |
text |
preview |
Benchmark results |
Leaderboards |
Coming soon |
Compare published results across the UniqueShip dataset splits. Rankings, evaluation metrics, and submission guidance will be available following the dataset release. |
| columns |
rows |
Rank |
Submission |
Dataset split |
Score |
Updated |
|
4 |
|
|
| inquiries |
| subheadline |
headline |
text |
benefits |
form |
Inquiries |
Need additional information or a different split? |
Current dataset splits are available to <a href="#releases" class="text-primary hover:underline">download directly</a> — no request or approval is required. Use this form if you have questions, need additional information, or would like to request a split that is not currently available. |
Ask questions about the data or documentation |
Request additional or specialized dataset splits |
|
| action |
submit_label |
submitting_label |
error_message |
success |
fields |
|
Submit inquiry |
Submitting… |
We could not send your request. Check your connection and try again. |
| eyebrow |
headline |
text |
Request received |
Thanks for your submission |
We have received your request and will be in touch. |
|
| type |
name |
label |
autocomplete |
required |
width |
text |
entry.925008751 |
Full name |
name |
true |
half |
|
| type |
name |
label |
autocomplete |
required |
width |
email |
entry.2144715529 |
Work email |
email |
true |
half |
|
| type |
name |
label |
autocomplete |
required |
width |
text |
entry.241770130 |
Organization |
organization |
false |
full |
|
| type |
name |
label |
required |
width |
textarea |
entry.1262269016 |
Inquiry |
true |
full |
|
|
|
|
| citation |
| subheadline |
headline |
text |
code |
contact |
Citing this dataset |
Reference the dataset paper |
If you use UniqueShip, please cite the paper below. |
@inproceedings{hashemi_2026_uniqueship,
author = {Hashemi, Connor and Stout, Trevor and Hoogs, Anthony and Parham, Jason},
title = {UniqueShip: Mitigating Data Leakage in Acoustic Ship Classification Benchmark Datasets},
booktitle = {OCEANS 2026},
year = {2026},
pages = {TODO}
} |
| headline |
text |
button_label |
url |
Questions, corrections, or collaboration? |
Our team can help with access and research partnerships. |
Contact the team |
|
|
|