> ## Documentation Index
> Fetch the complete documentation index at: https://docs.edgeimpulse.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Model testing

> Evaluate a trained model on held-out test data, read the confusion matrix, and tune per-block thresholds.

When collecting data, we split the dataset into training and testing sets. The model was trained with only the training set, and the testing set is used to validate how well the model will perform on unseen data. This will ensure that the model has not learned to overfit the training data, which is a common occurrence.

Training accuracy tells you how well a model learned the data it was shown. Model testing shows how well it handles data it hasn't seen, which is closer to how it will perform on a device.

## Prerequisites

Make sure to have data samples on your test set, you can add data samples from the **Data Acquisition** page or the **Live Classification** page.

<Frame caption="Test dataset">
  <img src="https://mintcdn.com/edgeimpulse/b53JLZcI-O3Jueox/.assets/images/test-dataset.png?fit=max&auto=format&n=b53JLZcI-O3Jueox&q=85&s=c2c74e89129b7379a9c8872de3295f94" alt="Test dataset section listing samples reserved for model testing" width="1421" height="1000" data-path=".assets/images/test-dataset.png" />
</Frame>

Testing runs against the impulse that is currently selected, so a project with several [experiments](/studio/projects/experiments) needs each impulse tested separately before the results are comparable.

## Running a test

To test your model, go to **Model testing**, select the desired model version from the dropdown (either **Unoptimized (float32)** or **Quantized (int8)**), and click **Test all**. The model will classify all of the test set samples and give you an overall accuracy of how your model performed.

<Info>
  Quantized (int8) model is not enabled by default and the fist step of enabling is in the settings menu beside the **Classify all** button
</Info>

Running a test generates features for every test sample, runs them through the impulse, and compares each prediction against the label you gave the sample.

<Frame caption="classify all test images">
  <img src="https://mintcdn.com/edgeimpulse/lnCwBUvZVOz6Veyh/.assets/images/model-testing.png?fit=max&auto=format&n=lnCwBUvZVOz6Veyh&q=85&s=ef0bcbeced8bbab8c021a8c73296ad05" alt="Model testing page with the Classify all action for test images" width="1319" height="1000" data-path=".assets/images/model-testing.png" />
</Frame>

### Float32 vs int8 models

You can test your model using either the **float32** or **int8** quantized version. The **float32** version offers higher precision but may use more resources, while the **int8** quantized version is optimized for memory and computational efficiency, making it suitable for edge devices with limited resources.

Test both variants. Quantization maps weights and activations to 8-bit integers, which usually costs a little accuracy in exchange for a smaller, faster model. Results are stored per variant, so once you have tested both you can switch between them in the dropdown and compare. Test the variant you plan to deploy.

## Reading the results

### Accuracy

The headline accuracy is the share of test samples the model classified correctly: samples classified as their expected label, divided by the total number of samples that counted towards the result. Studio also reports accuracy per class. Check it as well, because a model can be 95% accurate overall but 40% accurate on a rare class.

Depending on the project type, additional metrics are reported alongside accuracy, such as a balanced accuracy score for imbalanced datasets, separate anomaly and no-anomaly scores for anomaly detection, an F1 score for object detection, and a mean squared error for regression.

### Confusion matrix

This is also accompanied by a confusion matrix to show you how your model performs for each class and an interactive feature explorer that lets you click on a sample to easily visualize this dedicated result.

<Frame caption="Model testing confusion matrix">
  <img src="https://mintcdn.com/edgeimpulse/lnCwBUvZVOz6Veyh/.assets/images/model-testing-results.png?fit=max&auto=format&n=lnCwBUvZVOz6Veyh&q=85&s=66f81b9ecd1ae406bdaeb1969ad0d3c6" alt="Confusion matrix summarizing model testing accuracy by class" width="710" height="1000" data-path=".assets/images/model-testing-results.png" />
</Frame>

Each row is the true label of the samples you provided, and each column is the label the model predicted for them. The diagonal shows the correct predictions, and everything off the diagonal is a mistake. A cluster of samples in one off-diagonal cell means the model confuses those two classes. This is often a data problem, such as too few examples or examples that look alike after processing, rather than a model architecture problem.

### The sample table

The model testing data table has some quick actions available for each samples:

<Frame caption="Model testing data table">
  <img src="https://mintcdn.com/edgeimpulse/lnCwBUvZVOz6Veyh/.assets/images/model-testing-table.png?fit=max&auto=format&n=lnCwBUvZVOz6Veyh&q=85&s=2926343867ed99ca4d1aaf0127c1a7fc" alt="Model testing table listing sample predictions and expected labels" width="1109" height="1000" data-path=".assets/images/model-testing-table.png" />
</Frame>

<Warning>
  #### Limitation for anomaly detection

  Make sure to label your samples exactly as `anomaly` or `no anomaly` in your test dataset so they can be used in the F1 score calculation for anomaly detection projects. We are working on making this more flexible.
</Warning>

Also note that the samples who does not match the known classes for the classifiers or `anomaly` for anomaly detection learning blocks are ignored from the accuracy or the F1 score calculation:

<Frame caption="Ignored samples">
  <img src="https://mintcdn.com/edgeimpulse/lnCwBUvZVOz6Veyh/.assets/images/model-testing-ignored-samples.png?fit=max&auto=format&n=lnCwBUvZVOz6Veyh&q=85&s=27b67873aa34d3bc382c2645ff033213" alt="Ignored samples dialog listing test samples excluded from results" width="1396" height="750" data-path=".assets/images/model-testing-ignored-samples.png" />
</Frame>

## Setting confidence threshold

Every learning block has a threshold. This can be the minimum confidence that a neural network needs to have, or the maximum anomaly score before a sample is tagged as an anomaly. You can configure these thresholds to tweak the sensitivity of these learning blocks. This affects both live classification and model testing.

<Frame caption="Setting confidence threshold">
  <img src="https://mintcdn.com/edgeimpulse/sFtdfPhSpbLZ2-cz/.assets/images/studio-model-testing-gmm-set-confidence-threshold.png?fit=max&auto=format&n=sFtdfPhSpbLZ2-cz&q=85&s=07136885b0a90cde855c2e1866eedc93" alt="Confidence threshold control for setting anomaly detection cutoffs" width="1468" height="730" data-path=".assets/images/studio-model-testing-gmm-set-confidence-threshold.png" />
</Frame>

<br />

<Frame caption="Setting confidence threshold values">
  <img src="https://mintcdn.com/edgeimpulse/lnCwBUvZVOz6Veyh/.assets/images/model-testing-confidence-threshold.png?fit=max&auto=format&n=lnCwBUvZVOz6Veyh&q=85&s=3fbd8a406557b40989238665b3f8d688" alt="Confidence threshold dialog with editable class threshold values" width="1600" height="810" data-path=".assets/images/model-testing-confidence-threshold.png" />
</Frame>

Each threshold has a name and help text describing what it controls, and some have a suggested value Edge Impulse derived from your trained model. A threshold doesn't change the model, only where the cut-off sits between acting on a prediction and discarding it. Raising it means fewer false positives and more missed detections; lowering it does the reverse. Pick the direction that matches the cost of each mistake in your application.

Changing a threshold doesn't require retraining or regenerating features. Studio recalculates the existing test results against the new threshold.

## Evaluating individual samples

To see a classification in detail, go to the sample you want to evaluate, click the three dots next to it, and select **Show classification**. A new window shows the expected outcome and the model's prediction with its accuracy, which can help you see why an item was misclassified.

<Frame caption="Classification result. Showing the conclusions, the raw data and processed features in one overview.">
  <img src="https://mintcdn.com/edgeimpulse/MguyU0DkkpALWBaW/.assets/images/tutorial-continuous-motion-live-classification.png?fit=max&auto=format&n=MguyU0DkkpALWBaW&q=85&s=1eb20b49d42d027f481a8c8bcf807d11" alt="Classification result page with conclusions, raw data, and processed features" width="675" height="1000" data-path=".assets/images/tutorial-continuous-motion-live-classification.png" />
</Frame>

## Testing programmatically

| Task | Endpoint |
| - | - |
| Run a test against the test dataset, for one or more model variants | [Classify](/apis/studio/jobs/classify) |
| Set per-block thresholds, such as a minimum confidence or maximum anomaly score | [Set thresholds](/apis/studio/impulse/set-thresholds) |
| Recalculate test results after a threshold change, without regenerating features | [Regenerate model testing summary](/apis/studio/impulse/regenerate-model-testing-summary) |

Like the rest of the impulse endpoints, these take an optional `impulseId`, so you can test each experiment in a project from a script and collect the results together.

## Additional resources

* [Experiments](/studio/projects/experiments)
* [Live classification](/studio/projects/live-classification)
* [Performance calibration](/studio/projects/performance-calibration)
* [Increasing model performance](/knowledge/guides/increasing-model-performance)
