polars-avro project¶
polars_avro module¶
Polars io plugin for reading and writing Apache Avro files.
Provides scan_avro, read_avro, write_avro, and AvroWriter. Most polars
types write directly (narrow ints widen, Time truncates to microseconds); only
Categorical, Enum, and out-of-range UInt64 values must be cast first. When
reading, the utf8_view option controls how UUIDs and nullable strings are
decoded — see scan_avro for details.
- exception polars_avro.AvroError¶
Bases:
Exception
- exception polars_avro.AvroSpecError¶
Bases:
ValueError
- class polars_avro.AvroWriter(dest: str | Path | BinaryIO, *, schema: Schema | None = None, codec: Codec | None = None)¶
Bases:
objectIncrementally write DataFrames to an Avro file.
Most polars types write directly (narrow ints widen, Time truncates to microseconds); only Categorical, Enum, and out-of-range UInt64 values must be cast before writing — see the README for workarounds.
- close() None¶
- write(batch: DataFrame) None¶
- class polars_avro.Codec¶
Bases:
object- Bzip2 = Codec.Bzip2¶
- Deflate = Codec.Deflate¶
- Snappy = Codec.Snappy¶
- Xz = Codec.Xz¶
- Zstandard = Codec.Zstandard¶
- exception polars_avro.EmptySources¶
Bases:
ValueError
- polars_avro.read_avro(sources: Sequence[str | Path | BinaryIO] | str | Path | BinaryIO, *, columns: Sequence[int | str] | None = None, n_rows: int | None = None, row_index_name: str | None = None, row_index_offset: int = 0, rechunk: bool = False, batch_size: int = 32768, glob: bool = True, strict: bool = False, utf8_view: bool = False, storage_options: Mapping[str, str] | None = None) DataFrame¶
Read an Avro file into a DataFrame.
- Parameters:
sources (The source(s) to scan: local file paths, cloud URLs (
s3://,) –gs://,az://,http(s)://, …), or readable binary buffers. Binary buffers must be seekable (supportseek/tell); the reader rewinds them to read headers and to rewind for projection. Cloud URLs requirefsspec(plus the relevant backend, e.g.s3fs). Sources are read in the order given, so row order androw_indexfollow the argument order regardless of source kind.columns (The columns to select.)
n_rows (The number of rows to read.)
row_index_name (The name of the row index column, or None to not add one.)
row_index_offset (The offset to start the row index at.)
rechunk (Whether to rechunk the DataFrame after reading.)
batch_size (How many rows to attempt to read at a time.)
glob (Whether to use globbing to find files.)
strict (Whether to enable stricter avro union handling, rejecting unions) – where
nullis not the first branch ([T, "null"]rather than["null", T]) instead of accepting them.utf8_view (Whether to read strings as views. When
False(default),) – UUIDs are read as binary and nullable strings preserve nulls. WhenTrue, UUIDs are read as formatted strings and nulls in nullable strings are replaced with""(lossy). Since polars tends to work with string views internally,Trueis likely faster.storage_options (Extra options forwarded to
fsspec.openfor cloud URLs.)
- polars_avro.scan_avro(sources: Sequence[str | Path | BinaryIO] | str | Path | BinaryIO, *, batch_size: int = 1024, glob: bool = True, strict: bool = False, utf8_view: bool = False, storage_options: Mapping[str, str] | None = None) LazyFrame¶
Scan Avro files.
- Parameters:
sources (The source(s) to scan: local file paths, cloud URLs (
s3://,) –gs://,az://,http(s)://, …), or readable binary buffers. Binary buffers must be seekable (supportseek/tell); the reader rewinds them to read headers and to rewind for projection. Cloud URLs requirefsspec(plus the relevant backend, e.g.s3fs). Sources are read in the order given, so row order androw_indexfollow the argument order regardless of source kind.batch_size (How many rows to attempt to read at a time.)
glob (Whether to use globbing to find files (local paths only).)
strict (Whether to enable stricter avro union handling, rejecting unions) – where
nullis not the first branch ([T, "null"]rather than["null", T]) instead of accepting them.utf8_view (Whether to read strings as views. When
False(default),) – UUIDs are read as binary and nullable strings preserve nulls. WhenTrue, UUIDs are read as formatted strings and nulls in nullable strings are replaced with""(lossy). Since polars tends to work with string views internally,Trueis likely faster.storage_options (Extra options forwarded to
fsspec.openfor cloud URLs.)
- polars_avro.write_avro(batches: DataFrame | Iterable[DataFrame], dest: str | Path | BinaryIO, *, schema: Schema | None = None, codec: Codec | None = None) None¶
Write a DataFrame or iterable of DataFrames to an Avro file.
Most polars types write directly (narrow ints widen, Time truncates to microseconds); only Categorical, Enum, and out-of-range UInt64 values must be cast before writing — see the README for workarounds.
- Parameters:
batches (A DataFrame or iterable of DataFrames to write.)
dest (The file path or writable binary buffer to write to.)
schema (The schema to use. If None, inferred from the first batch.)
codec (The compression codec to use, or None for no compression.)