Write this DataFrame to one or more (Geo)Parquet files. For input that contains geometry columns, GeoParquet metadata is written such that suitable readers can recreate Geometry/Geography types when reading the output and potentially read fewer row groups when only a subset of the file is needed for a given query.

sd_write_parquet(
  .data,
  path,
  options = NULL,
  partition_by = character(0),
  sort_by = character(0),
  single_file_output = NULL,
  geoparquet_version = "1.0",
  overwrite_bbox_columns = FALSE,
  max_row_group_size = NULL,
  compression = NULL
)

Arguments

.data

A sedonadb_dataframe or an object that can be coerced to one.

path

A filename or directory to which parquet file(s) should be written

options

A named list of key/value options to be used when constructing a parquet writer. Common options are exposed as other arguments to sd_write_parquet(); however, this argument allows setting any DataFusion Parquet writer option. If an option is specified here and by another argument to this function, the value specified as an explicit argument takes precedence.

partition_by

A character vector of column names to partition by. If non-empty, applies hive-style partitioning to the output

sort_by

A character vector of column names to sort by. Currently only ascending sort is supported

single_file_output

Use TRUE or FALSE to force writing a single Parquet file vs. writing one file per partition to a directory. By default, a single file is written if partition_by is unspecified and path ends with .parquet

geoparquet_version

GeoParquet metadata version to write if output contains one or more geometry columns. The default ("1.0") is the most widely supported and will result in geometry columns being recognized in many readers; however, only includes statistics at the file level. Use "1.1" to compute an additional bounding box column for every geometry column in the output: some readers can use these columns to prune row groups when files contain an effective spatial ordering. The extra columns will appear just before their geometry column and will be named "[geom_col_name]_bbox" for all geometry columns except "geometry", whose bounding box column name is just "bbox"

overwrite_bbox_columns

Use TRUE to overwrite any bounding box columns that already exist in the input. This is useful in a read -> modify -> write scenario to ensure these columns are up-to-date. If FALSE (the default), an error will be raised if a bbox column already exists

max_row_group_size

Target maximum number of rows in each row group. Defaults to the global configuration value (1M rows).

compression

Sets the Parquet compression codec. Valid values are: uncompressed, snappy, gzip(level), brotli(level), lz4, zstd(level), and lz4_raw. Defaults to the global configuration value (zstd(3)).

Value

The input, invisibly

Examples

tmp_parquet <- tempfile(fileext = ".parquet")

sd_sql("SELECT ST_Point(1, 2, 4326) as geom") |>
  sd_write_parquet(tmp_parquet)

sd_read_parquet(tmp_parquet)
#> <sedonab_dataframe: ?? x 1>
#> ┌────────────┐
#> │    geom    │
#> │  geometry  │
#> ╞════════════╡
#> │ POINT(1 2) │
#> └────────────┘
#> Preview of up to 6 row(s)
unlink(tmp_parquet)