Sequence dataset format

A sequence dataset exports as three Parquet files. The viam dataset export command writes them to a zip, and a custom training job receives them as command-line arguments. Both use the same schema. To produce an export you can inspect, follow the sequences tutorial.

Export layout

Running viam dataset export on a sequence dataset writes:

<destination>/
  <dataset-id>.zip
    binary_data.parquet
    tabular_data.parquet
    sequences.parquet
  binary_data/
    <binary-data-id><extension>

Binary data IDs contain slashes, so the image files sit in nested folders under binary_data/.

In an exported binary_data.parquet, the path column is relative to the export directory. In a training job, path is an absolute path, starting with /gcs/, that your script can open directly.

binary_data.parquet

One row per image in each sequence.

ColumnTypeDescription
sequence_idstringID of the sequence the image belongs to.
timestamptimestamp (microseconds)When the machine captured the image.
part_idstringID of the machine part that captured the image.
component_namestringName of the resource that produced the image.
method_namestringMethod that produced the image, for example GetImages.
pathstringLocation of the image file. See Export layout.
classificationslist of {label, confidence}Classification annotations on the image.
bounding_boxeslist of {label, confidence, x_min, y_min, x_max, y_max}Bounding box annotations, with coordinates normalized to the image size.

tabular_data.parquet

One row per reading in each sequence.

ColumnTypeDescription
sequence_idstringID of the sequence the reading belongs to.
timestamptimestamp (microseconds)When the machine captured the reading.
part_idstringID of the machine part that captured the reading.
component_namestringName of the resource that produced the reading.
method_namestringMethod that produced the reading, for example Readings.
payloadstringThe method’s response, as a JSON string. For a sensor’s Readings, for example: {"readings":{"a":1.0,"b":2.0}}.

sequences.parquet

One row per sequence.

ColumnTypeDescription
sequence_idstringID of the sequence.
tagslist of stringThe sequence’s tags.
start_attimestamp (microseconds)Start of the sequence’s time window.
end_attimestamp (microseconds)End of the sequence’s time window.

The API and SDKs call these fields start_time and end_time.

Join the files

Join binary_data.parquet and tabular_data.parquet to sequences.parquet on sequence_id.

The export doesn’t resample or align the data. Each resource keeps its own sample rate, so a script that needs images and readings at the same moments has to match them itself, for example by nearest timestamp. All timestamps are in UTC, and libraries such as pandas load them as timezone-aware UTC values.

Training script arguments

A custom training job on a sequence dataset passes your script the path to each Parquet file:

ArgumentDescription
--binary_data_filePath to binary_data.parquet.
--tabular_data_filePath to tabular_data.parquet.
--sequences_filePath to sequences.parquet.
--model_output_directoryWhere to write the trained model. Viam sets it.

A job on a sequence dataset doesn’t get --dataset_file. A job on a binary dataset gets --dataset_file and none of the three Parquet file arguments. Your own arguments can’t reuse --dataset_file, --model_output_directory, or the three sequence file names.

See Custom training scripts.