Skip to main content
Only available on the Enterprise planThis feature is only available on the Enterprise plan. Review our plans and pricing or sign up for our free expert-led trial today.
Organization data holds datasets that you can use across projects. A project’s dataset is scoped to that project and stored by Edge Impulse; an organization dataset lives in your own cloud storage and can be queried, transformed, validated, and imported into as many projects as you like. With more than one project, this saves uploading the same files to each project and keeping the copies in sync.

How organization data is structured

Organization data has three layers:
  • A bucket is your cloud storage: S3, Google Cloud Storage, Azure Blob Storage, or any S3-compatible service. Edge Impulse connects to it with credentials you supply.
  • A dataset is a pointer to a path inside that bucket, plus the metadata Edge Impulse keeps about it. Creating a dataset does not copy or move your files. Your bucket stays the source of truth, and files added to that path by any other tool show up in Edge Impulse.
  • A project import copies files from the dataset into a project, converting them to a format the ingestion service accepts.
Deleting a dataset in Edge Impulse doesn’t delete your files, and you can change files directly in the bucket.
You can also create two datasets over the same bucket path, one generic and one clinical, to get easy uploads from the first and query support from the second.

Health reference design

We have built a health reference design that describes an end-to-end ML workflow for building a wearable health product using Edge Impulse. In this reference resign, we want to help you understand how to create a full clinical data pipeline by using a public dataset from the PPG-DaLiA repository. This tutorial will guide you through the following steps:

Buckets

Before we get started, you must link your organization with one or more storage buckets. Further details about how to integrate with cloud storage providers can be found in the Cloud data storage document.

Datasets

Two types of dataset structures can be used - Generic datasets (default) and Clinical datasets.
There is no required format for data files. You can upload data in a wide range of formats, whether it’s CSV, Parquet, or a proprietary data format.However, to import data items to an Edge Impulse project, you will need to use the right format as our studio ingestion API only supports these formats:
  • JPG, PNG images
  • MP4, AVI video files
  • WAV audio files
  • JSON/CBOR files in the Edge Impulse data acquisition format
  • CSV files
Tip: You can use transformation blocks to convert your data
Organization datasets overview listing connected datasets and storage information

Datasets overview

The default dataset structure is a file-based one, no matter the directory structure:For example:
or:
Note that you will be able to associate the labels of your data items from the file name or the directory name when importing your data in a project.

Create a new dataset

Once you successfully linked your storage bucket to your organization, head to the Datasets tab and click on + Add new dataset:
Add new dataset dialog with dataset name and type fields

Add new dataset

Fill out the following form:
Add dataset form showing storage bucket and dataset path settings

Add dataset

Click on Create dataset

Data

With your datasets imported, you can now navigate into your dataset, create folders, query your dataset, add data items and import your data to an Edge Impulse project.
Default view
The default view lets you navigate in your bucket following the directory structure. You can create folders with the + New folder button. To add data, drag and drop files and folders onto the right panel, and they upload to your bucket automatically.
Data items overview for browsing files in an organization dataset

Data items overview

Adding data to your project

Go to the Actions…->Import data into a project, select the project you wish to import to and click Next, Configure how to label this data:
Upload files dialog for adding organization data to a project

Uploading Files

This will import the data into the project and optionally create a new label for each file in the dataset. This labeling step helps you keep track of different classes or categories within your data. After importing the data into the project, in the Next, post-sync actions step, you can configure a data pipeline to automatically retrieve and trigger actions in your project:
Post-sync labeling action dialog for assigning labels to imported files

Label your files

Previewing data

We also have added a data preview feature, allowing you to visualize certain types of data directly within the organization data tab. Supported data types include tables (CSV/Parquet), images, PDFs, audio files (WAV/MP3), and text files (TXT/JSON). This feature gives you a quick overview of your data and helps ensure its integrity and correctness.
CSV and Parquet data visualization table in the data items overview

Data items overview - CSV/Parquet type


Image data visualization grid in the data items overview

Data items overview - image type

API reference

Additional resources

Any questions, or interested in the enterprise version of Edge Impulse? Contact us for more information.