Skip to main content
The Advanced settings on the Dataset tab control how your project uses its data. For example, an explicit validation set changes how your project divides data between training, validation, and testing.
Advanced settings modal with explicit validation set and reset dataset options

Advanced settings modal

To access the advanced settings, go to your project and navigate to the data acquisition page. On the dataset tab, click the ⋮ in the top right of the dataset pane, and then click on Advanced settings menu item. This will open a modal with the settings available.
Dataset pane menu with the Advanced settings action highlighted

Advanced settings menu item

Explicit validation set

Why a validation set exists

Your data splits into three sets. The training set is what the model learns from. The validation set is held back and scored after every epoch, so you can watch for overfitting while training. The test set is used only once training is done, by model testing, to estimate real-world performance. Validation and test stay separate because you make decisions based on the validation score, such as when to stop training. Data you’ve made decisions on is no longer unseen, which is why the test set stays untouched until the end.

What the setting changes

By default, Edge Impulse automatically splits your training data into the training and validation sets. Enabling the explicit validation set setting allows you to specify the percentage of data to use for your validation set, explore what samples are contained in the set, and move samples to/from the validation set from/to your training or test sets.
Validation set tab showing samples after the explicit validation set is enabled

Validation set samples after enabling explicit validation set

This gives you more control over the data used for the validation set, such as ensuring it contains specific samples or has a certain distribution. It also allows you to have a fixed validation set that does not change as you add new data to your training set. After enabling the explicit validation set setting, a field will appear in the Define dataset split modal to specify the percentage of data to use for the validation set. If you subsequently disable the explicit validation set setting, any samples in your validation set will be moved back to your training set and the validation split percentage field will be removed from the define dataset split modal. Future training jobs will use the standard validation split behaviour.

When to enable it

Turn this on when the validation set’s composition affects whether you can trust the score:
  • Data is grouped by subject, device, or session, and a group must be kept out of training entirely to avoid leakage. Pair this with the grouping controls in Define dataset split.
  • You are comparing experiments and want every impulse validated against the same samples.
  • Your dataset is small or imbalanced, so an automatically drawn validation set may not contain enough examples of a rare class.
The setting is exposed in the API as explicitValidationEnabled on Update project.
Warning dialog confirming that disabling the explicit validation set will reset splits

Disable explicit validation set warning

Additional resources