Sequence dataset format
A sequence dataset exports as three Parquet files.
The viam dataset export command writes them to a zip, and a custom training job receives them as command-line arguments.
Both use the same schema.
To produce an export you can inspect, follow the sequences tutorial.
Export layout
Running viam dataset export on a sequence dataset writes:
<destination>/
<dataset-id>.zip
binary_data.parquet
tabular_data.parquet
sequences.parquet
binary_data/
<binary-data-id><extension>
Binary data IDs contain slashes, so the image files sit in nested folders under binary_data/.
In an exported binary_data.parquet, the path column is relative to the export directory.
In a training job, path is an absolute path, starting with /gcs/, that your script can open directly.
binary_data.parquet
One row per image in each sequence.
| Column | Type | Description |
|---|---|---|
sequence_id | string | ID of the sequence the image belongs to. |
timestamp | timestamp (microseconds) | When the machine captured the image. |
part_id | string | ID of the machine part that captured the image. |
component_name | string | Name of the resource that produced the image. |
method_name | string | Method that produced the image, for example GetImages. |
path | string | Location of the image file. See Export layout. |
classifications | list of {label, confidence} | Classification annotations on the image. |
bounding_boxes | list of {label, confidence, x_min, y_min, x_max, y_max} | Bounding box annotations, with coordinates normalized to the image size. |
tabular_data.parquet
One row per reading in each sequence.
| Column | Type | Description |
|---|---|---|
sequence_id | string | ID of the sequence the reading belongs to. |
timestamp | timestamp (microseconds) | When the machine captured the reading. |
part_id | string | ID of the machine part that captured the reading. |
component_name | string | Name of the resource that produced the reading. |
method_name | string | Method that produced the reading, for example Readings. |
payload | string | The method’s response, as a JSON string. For a sensor’s Readings, for example: {"readings":{"a":1.0,"b":2.0}}. |
sequences.parquet
One row per sequence.
| Column | Type | Description |
|---|---|---|
sequence_id | string | ID of the sequence. |
tags | list of string | The sequence’s tags. |
start_at | timestamp (microseconds) | Start of the sequence’s time window. |
end_at | timestamp (microseconds) | End of the sequence’s time window. |
The API and SDKs call these fields start_time and end_time.
Join the files
Join binary_data.parquet and tabular_data.parquet to sequences.parquet on sequence_id.
The export doesn’t resample or align the data. Each resource keeps its own sample rate, so a script that needs images and readings at the same moments has to match them itself, for example by nearest timestamp. All timestamps are in UTC, and libraries such as pandas load them as timezone-aware UTC values.
Training script arguments
A custom training job on a sequence dataset passes your script the path to each Parquet file:
| Argument | Description |
|---|---|
--binary_data_file | Path to binary_data.parquet. |
--tabular_data_file | Path to tabular_data.parquet. |
--sequences_file | Path to sequences.parquet. |
--model_output_directory | Where to write the trained model. Viam sets it. |
A job on a sequence dataset doesn’t get --dataset_file.
A job on a binary dataset gets --dataset_file and none of the three Parquet file arguments.
Your own arguments can’t reuse --dataset_file, --model_output_directory, or the three sequence file names.
Was this page helpful?
Glad to hear it! If you have any other feedback please let us know:
We're sorry about that. To help us improve, please tell us what we can do better:
Thank you!