Skip to main content
When collecting data, we split the dataset into training and testing sets. The model was trained with only the training set, and the testing set is used to validate how well the model will perform on unseen data. This will ensure that the model has not learned to overfit the training data, which is a common occurrence. Training accuracy tells you how well a model learned the data it was shown. Model testing shows how well it handles data it hasn’t seen, which is closer to how it will perform on a device.

Prerequisites

Make sure to have data samples on your test set, you can add data samples from the Data Acquisition page or the Live Classification page.
Test dataset section listing samples reserved for model testing

Test dataset

Testing runs against the impulse that is currently selected, so a project with several experiments needs each impulse tested separately before the results are comparable.

Running a test

To test your model, go to Model testing, select the desired model version from the dropdown (either Unoptimized (float32) or Quantized (int8)), and click Test all. The model will classify all of the test set samples and give you an overall accuracy of how your model performed.
Quantized (int8) model is not enabled by default and the fist step of enabling is in the settings menu beside the Classify all button
Running a test generates features for every test sample, runs them through the impulse, and compares each prediction against the label you gave the sample.
Model testing page with the Classify all action for test images

classify all test images

Float32 vs int8 models

You can test your model using either the float32 or int8 quantized version. The float32 version offers higher precision but may use more resources, while the int8 quantized version is optimized for memory and computational efficiency, making it suitable for edge devices with limited resources. Test both variants. Quantization maps weights and activations to 8-bit integers, which usually costs a little accuracy in exchange for a smaller, faster model. Results are stored per variant, so once you have tested both you can switch between them in the dropdown and compare. Test the variant you plan to deploy.

Reading the results

Accuracy

The headline accuracy is the share of test samples the model classified correctly: samples classified as their expected label, divided by the total number of samples that counted towards the result. Studio also reports accuracy per class. Check it as well, because a model can be 95% accurate overall but 40% accurate on a rare class. Depending on the project type, additional metrics are reported alongside accuracy, such as a balanced accuracy score for imbalanced datasets, separate anomaly and no-anomaly scores for anomaly detection, an F1 score for object detection, and a mean squared error for regression.

Confusion matrix

This is also accompanied by a confusion matrix to show you how your model performs for each class and an interactive feature explorer that lets you click on a sample to easily visualize this dedicated result.
Confusion matrix summarizing model testing accuracy by class

Model testing confusion matrix

Each row is the true label of the samples you provided, and each column is the label the model predicted for them. The diagonal shows the correct predictions, and everything off the diagonal is a mistake. A cluster of samples in one off-diagonal cell means the model confuses those two classes. This is often a data problem, such as too few examples or examples that look alike after processing, rather than a model architecture problem.

The sample table

The model testing data table has some quick actions available for each samples:
Model testing table listing sample predictions and expected labels

Model testing data table

Limitation for anomaly detection

Make sure to label your samples exactly as anomaly or no anomaly in your test dataset so they can be used in the F1 score calculation for anomaly detection projects. We are working on making this more flexible.
Also note that the samples who does not match the known classes for the classifiers or anomaly for anomaly detection learning blocks are ignored from the accuracy or the F1 score calculation:
Ignored samples dialog listing test samples excluded from results

Ignored samples

Setting confidence threshold

Every learning block has a threshold. This can be the minimum confidence that a neural network needs to have, or the maximum anomaly score before a sample is tagged as an anomaly. You can configure these thresholds to tweak the sensitivity of these learning blocks. This affects both live classification and model testing.
Confidence threshold control for setting anomaly detection cutoffs

Setting confidence threshold


Confidence threshold dialog with editable class threshold values

Setting confidence threshold values

Each threshold has a name and help text describing what it controls, and some have a suggested value Edge Impulse derived from your trained model. A threshold doesn’t change the model, only where the cut-off sits between acting on a prediction and discarding it. Raising it means fewer false positives and more missed detections; lowering it does the reverse. Pick the direction that matches the cost of each mistake in your application. Changing a threshold doesn’t require retraining or regenerating features. Studio recalculates the existing test results against the new threshold.

Evaluating individual samples

To see a classification in detail, go to the sample you want to evaluate, click the three dots next to it, and select Show classification. A new window shows the expected outcome and the model’s prediction with its accuracy, which can help you see why an item was misclassified.
Classification result page with conclusions, raw data, and processed features

Classification result. Showing the conclusions, the raw data and processed features in one overview.

Testing programmatically

Like the rest of the impulse endpoints, these take an optional impulseId, so you can test each experiment in a project from a script and collect the results together.

Additional resources