module pixeltable.io
Functions for importing and exporting Pixeltable data.func export_csv()
Signature
- String, Int, Float, Bool: native CSV representation
- Timestamp, Date: ISO 8601 string representation
- UUID: string representation
- Json: JSON-encoded string
- Array: JSON-encoded string (via
tolist()) - Binary: excluded from export (not representable in CSV)
- Image, Video, Audio, Document: file path or URL string
table_or_query(pxt.Table | pxt.Query): Table or Query to export.file_path(str | Path): Path to the output CSV file.delimiter(str, default:','): Field delimiter character. Default','.quoting(int, default:0): CSV quoting style (acsv.QUOTE_*constant). Defaultcsv.QUOTE_MINIMAL.
func export_iceberg()
Signature
RecordBatches, the size of
which can be controlled with the batch_size_bytes parameter.
Requirements:
pip install pyiceberg
table_or_query(pxt.Table | pxt.Query): PixeltableTableorQueryto export.catalog(Catalog): An IcebergCataloginstance to write the table into.table_name(str): Fully-qualified Iceberg table identifier (e.g.'my_namespace.my_table'). If the namespace does not exist, it will be created.batch_size_bytes(int, default:134217728): Maximum size in bytes for each in-memory pyarrow batch.if_exists(Literal['error', 'replace', 'append'], default:'error'): Determines the behavior if the table already exists. Must be one of the following:'error': raise an error'replace': drop the existing table and create a new one'append': append to the existing table (source schema must be compatible)
schema_overrides(Mapping[str, pa.DataType] | None): If specified, then for each (name, type) pair inschema_overrides, the column with namenamewill be given typetype, instead of being inferred from the Pixeltable schema. The keys inschema_overridesshould be the column names of the Pixeltable schema.
func export_images_as_fo_dataset()
Signature
frame_iterator) will first be written to disk in the specified
image_format.
The label parameters accept one or more sets of labels of each type. If a single Expr is provided, then it will
be exported as a single set of labels with a default name such as classifications.
(The single set of labels may still containing multiple individual labels; see below.)
If a list of Exprs is provided, then each one will be exported as a separate set of labels with a default name
such as classifications, classifications_1, etc. If a dictionary of Exprs is provided, then each entry will
be exported as a set of labels with the specified name.
Requirements:
pip install fiftyone
-
tbl(pxt.Table): The table from which to export data. -
images(exprs.Expr): A column or expression that contains the images to export. -
image_format(str, default:'webp'): The format to use when writing out images for export. -
classifications(exprs.Expr | list[exprs.Expr] | dict[str, exprs.Expr] | None): Optional image classification labels. If a singleExpris provided, it must be a table column or an expression that evaluates to a list of dictionaries. Each dictionary in the list corresponds to an image class and must have the following structure:If multipleExprs are provided, each one must evaluate to a list of such dictionaries. -
detections(exprs.Expr | list[exprs.Expr] | dict[str, exprs.Expr] | None): Optional image detection labels. If a singleExpris provided, it must be a table column or an expression that evaluates to a list of dictionaries. Each dictionary in the list corresponds to an image detection, and must have the following structure:If multipleExprs are provided, each one must evaluate to a list of such dictionaries.
'fo.Dataset': A Voxel51 dataset.
image column of the table tbl as a Voxel51 dataset, using classification
labels from tbl.classifications:
func export_json()
Signature
- String: string
- Int: number
- Float: number
- Bool: boolean
- Timestamp: ISO 8601 string
- Date: ISO 8601 string
- UUID: string
- Json: native JSON value (object, array, etc.)
- Array: nested JSON array (via
tolist()) - Binary: excluded from export (not representable in JSON)
- Image, Video, Audio, Document: file path or URL string
table_or_query(pxt.Table | pxt.Query): Table or Query to export.file_path(str | Path): Path to the output JSONL file.
func export_lancedb()
Signature
RecordBatches, the size of which can be controlled with the batch_size_bytes parameter.
Requirements:
pip install lancedb pylance
table_or_query(Any): Table or Query to export.db_uri(Path): Local Path to the LanceDB database.table_name(Any): Name of the table in the LanceDB database.batch_size_bytes(Any): Maximum size in bytes for each batch.if_exists(Literal['error', 'overwrite', 'append'], default:'error'): Determines the behavior if the table already exists. Must be one of the following:'error': raise an error'overwrite': overwrite the existing table'append': append to the existing table
func export_parquet()
Signature
- String: string
- Int: int64
- Float: float32
- Bool: bool
- Timestamp: timestamp[us, tz=UTC]
- Date: date32
- UUID: uuid
- Binary: binary
-
Image: binary (when
inline_images=True) - Audio, Video, Document: string (file paths)
-
Array (requires shape to be known):
- fixed_shape_tensor for fixed-shape arrays
- list for ragged arrays (one or more dimensions are None)
-
Json: struct
- Schema is inferred from data via
pyarrow.infer_type() - Fields that contain empty dicts cannot be mapped to a Parquet type and will result in an exception
- Schema is inferred from data via
table_or_query(Any): Table or Query to export.parquet_path(Any): Path to directory to write the parquet files to.partition_size_bytes(Any): The maximum target size for each chunk. Default 100_000_000 bytes.inline_images(Any): If True, images are stored inline in the parquet file. This is useful for small images, to be imported as pytorch dataset. But can be inefficient for large images, and cannot be imported into pixeltable. If False, will raise an error if the Query has any image column. Default False.
func export_sql()
Signature
table_or_query(Any): Table or Query to export.target_table_name(Any): Name of the target table.db_connect_str(Any): Connection string to the target database.target_schema_name(Any): Optional name of the target schema.if_exists(Any): What to do if the target table already exists.- ‘error’: raise an error
- ‘replace’: drop the existing table and create a new one
- ‘insert’: insert new rows into the existing table
if_not_exists(Any): What to do if the target table does not exist.- ‘error’: raise an error
- ‘create’: create the table from the source schema
func import_csv()
Signature
import_pandas(table_path, pd.read_csv(filepath_or_buffer, **kwargs), schema=schema).
See the Pandas documentation for read_csv
for more details.
Returns:
pxt.Table: A handle to the newly createdTable.
func import_excel()
Signature
import_pandas(table_path, pd.read_excel(io, *args, **kwargs), schema=schema).
See the Pandas documentation for read_excel
for more details.
Returns:
pxt.Table: A handle to the newly createdTable.
func import_huggingface_dataset()
Signature
datasets library to be installed.
HuggingFace feature types are mapped to Pixeltable column types as follows:
Value(bool):Bool
Value(int*/uint*):Int
Value(float*):Float
Value(string/large_string):String
Value(timestamp*):Timestamp
Value(date*):DateClassLabel:String(converted to label names)Sequence/LargeListof numeric types:ArraySequence/LargeListof string:JsonSequence/LargeListof dicts:JsonArray2D-Array5D:Array(preserves shape)Image:ImageAudio:AudioVideo:VideoTranslation/TranslationVariableLanguages:Json
table_path(str): Path to the table.dataset(datasets.Dataset | datasets.DatasetDict | datasets.IterableDataset | datasets.IterableDatasetDict): An instance of any of the Huggingface dataset classes:datasets.Dataset,datasets.DatasetDict,datasets.IterableDataset,datasets.IterableDatasetDictschema_overrides(dict[str, Any] | None): If specified, then for each (name, type) pair inschema_overrides, the column with namenamewill be given typetype, instead of being inferred from theDatasetorDatasetDict. The keys inschema_overridesshould be the column names of theDatasetorDatasetDict(whether or not they are valid Pixeltable identifiers).primary_key(str | list[str] | None): The primary key of the table (seecreate_table()).kwargs(Any): Additional arguments to pass tocreate_table. An argument ofcolumn_name_for_splitmust be provided if the source is a DatasetDict. This column name will contain the split information. If None, no split information will be stored.
pxt.Table: A handle to the newly createdTable.
func import_json()
Signature
create_table() with
pxt.create_table(tbl_path, source=filepath_or_url, extra_args=kwargs, ...).
The contents of filepath_or_url are read and parsed as JSON internally (using json.loads(**kwargs)).
Parameters:
tbl_path(str): The name of the table to create.filepath_or_url(str): The path or URL of the JSON file.schema_overrides(dict[str, Any] | None): If specified, then columns inschema_overrideswill be given the specified types (seeimport_rows()).primary_key(str | list[str] | None): The primary key of the table (seecreate_table()).comment(str | None): A comment to attach to the table (seecreate_table()).kwargs(Any): Additional keyword arguments to pass tojson.loads.
pxt.Table: A handle to the newly createdTable.
func import_pandas()
Signature
DataFrame, with the
specified name. The schema of the table will be inferred from the DataFrame.
The column names of the new table will be identical to those in the DataFrame, as long as they are valid
Pixeltable identifiers. If a column name is not a valid Pixeltable identifier, it will be normalized according to
the following procedure:
- first replace any non-alphanumeric characters with underscores;
- then, preface the result with the letter ‘c’ if it begins with a number or an underscore;
- then, if there are any duplicate column names, suffix the duplicates with ‘_2’, ‘_3’, etc., in column order.
tbl_name(str): The name of the table to create.df(pd.DataFrame): The PandasDataFrame.schema_overrides(dict[str, Any] | None): If specified, then for each (name, type) pair inschema_overrides, the column with namenamewill be given typetype, instead of being inferred from theDataFrame. The keys inschema_overridesshould be the column names of theDataFrame(whether or not they are valid Pixeltable identifiers).
pxt.Table: A handle to the newly createdTable.
func import_parquet()
Signature
table(str): Fully qualified name of the table to import the data into.parquet_path(str): Path to an individual Parquet file or directory of Parquet files.schema_overrides(dict[str, Any] | None): If specified, then for each (name, type) pair inschema_overrides, the column with namenamewill be given typetype, instead of being inferred from the Parquet dataset. The keys inschema_overridesshould be the column names of the Parquet dataset (whether or not they are valid Pixeltable identifiers).primary_key(str | list[str] | None): The primary key of the table (seecreate_table()).kwargs(Any): Additional arguments to pass tocreate_table.
pxt.Table: A handle to the newly created table.
func import_rows()
Signature
{column_name: value, ...}. Pixeltable will attempt to infer the schema of the table from the
supplied data, using the most specific type that can represent all the values in a column.
If schema_overrides is specified, then for each entry (column_name, type) in schema_overrides,
Pixeltable will force the specified column to the specified type (and will not attempt any type inference
for that column).
All column types of the new table will be nullable unless explicitly specified as non-nullable in
schema_overrides.
Parameters:
tbl_path(str): The qualified name of the table to create.rows(list[dict[str, Any]]): The list of dictionaries to import.schema_overrides(dict[str, Any] | None): If specified, then columns inschema_overrideswill be given the specified types as described above.primary_key(str | list[str] | None): The primary key of the table (seecreate_table()).comment(str, default:''): A comment to attach to the table (seecreate_table()).
pxt.Table: A handle to the newly createdTable.
func import_sql()
Signature
-
selectable(sqlalchemy.sql.selectable.Selectable): A SQLAlchemySelectable(aTable, aselect()statement, ortext(...).columns(...)) describing the source rows. -
conn(sqlalchemy.engine.base.Engine | sqlalchemy.engine.base.Connection): A SQLAlchemyEngineorConnectionto executeselectableagainst. If aConnectionis passed, it must remain open and untouched (no commits, rollbacks, or other statements) for the duration of the import; rows are streamed directly from the server-side cursor. -
tbl_name(str): Pixeltable path of the destination table. -
schema_overrides(dict[str, typing.Any] | None): Optional per-column overrides applied on top of the inferred schema. Keys are column names; values accept any Pixeltable type spec recognized bypxt.create_table(eg,pxt.Image,pxt.Required[pxt.String]). -
primary_key(str | list[str] | None): An optional column name or list of column names to use as the primary key(s) of the table. Only applies when a new table is created; ignored when appending to an existing table. -
comment(str | None): An optional comment; its meaning is user-defined. Only applies when a new table is created; ignored when appending to an existing table. -
custom_metadata(typing.Any): Optional user-defined metadata to associate with the table. Must be a valid JSON-serializable object [str, int, float, bool, dict, list]. Only applies when a new table is created; ignored when appending to an existing table. -
if_exists(typing.Literal['error', 'append'], default:'error'): How to handle the destination table.'error': create the table; fail if it already exists.'append': append into the table if it already exists (verifying the source schema is compatible); otherwise create it.
-
on_error(typing.Literal['abort', 'ignore'], default:'abort'): Determines the behavior if an error occurs while evaluating a computed column or detecting an invalid media file (such as a corrupt image) for one of the inserted rows.- If
on_error='abort', then an exception will be raised and the rows will not be inserted. - If
on_error='ignore', then execution will continue and the rows will be inserted. Any cells with errors will have aNonevalue for that cell, with information about the error stored in the correspondingtbl.col_name.errortypeandtbl.col_name.errormsgfields.
- If
-
send_connect_url(bool, default:False): Determines where the query is run, and is only relevant when the destination table is hosted. It has no effect on a local table.- If
False, the query is run locally and the result is sent to the server. - If
True, the server connects to the source database itself (using the source connection URL, including credentials); this requires the server’s host to be able to reach the database and is faster for large imports.
- If
pixeltable.catalog.table.Table: The destinationTable.