EyeOnWater training dataset for assessing the inclusion of water images
Description
Data preprocessing In order to create a larger training dataset the set of original images (containing a total of 1700 images) are augmented, by rotating, displacing and resizing them. Using the following settings: Maximum rotation of 45 degrees in both directions Maximum displacement of 20% times the width or height Horizontal and vertical flip Maximum shear range of 20% times the width Pixel range of 10 units Data splitting The training dataset is 80% used for training, 10% for validation and 10% for prediction. Data labelling The dataset is divided into three classes, as previously mentioned. Initially, the training model was trained on just two classes: “water” and “nonWater.” However, it struggled to distinguish between images of acceptable water and those that did not meet the required standards. To address this, a third class, “water_bad,” was introduced. This class includes images of water that either show the ocean floor, or where a significant portion of the water is obscured by objects such as boats or docks. With the addition of the “water_bad” class, the model's ability to differentiate between acceptable and non-compliant water was improved. Parameters From the images the water quality can be obtained by comparing the water color to the 21 colors in the Forel-Ule scale. Parameter (P01): http://vocab.nerc.ac.uk/collection/P01/current/CLFORULE/ Parameter (P02) - see keywords: https://vocab.nerc.ac.uk/collection/P02/current/R410/ Data sources The images are taken by citizen scientists, often with a smartphone. Data quality To improve data quality and address class imbalance in our model, we applied data augmentation. The initial distribution of images was uneven, with 11,406 images in the “water_good” class, 1,789 in “water_bad,” and only 466 in the “other” class. To achieve a more balanced dataset, we augmented the images so that each class would contain approximately 3,000 images. This augmentation process enhanced the model’s ability to learn from each class more effectively, ensuring consistent representation across categories. Data resizing Larger images are resized to 256px by 256px, smaller images are excluded from the training dataset. Spatial coverage Images are taken on a global scale. Contact information For more information on the training dataset and/or the app, you can contact tjerk@maris.nl.
Resources
| Name |
Format |
Description |
Link |
|
0 |
|
http://data.europa.eu/88u/dataset/oai-zenodo-org-14017143 |
|
0 |
|
http://data.europa.eu/88u/dataset/oai-zenodo-org-14017143 |