Summary
Before decomposing pyiceberg/io/pyarrow.py into focused submodules (#3737, #3738), we should consolidate PyArrow-specific logic that currently lives outside the module. This ensures all PyArrow calls route through a single boundary, making the subsequent split clean and enabling future engine substitution.
Motivation
Per discussion in #3737, @rambleraptor noted that the first useful step is ensuring no PyArrow logic occurs outside pyarrow.py. Currently several modules import pyarrow directly and implement compute logic inline rather than delegating through pyiceberg.io.pyarrow.
When we later introduce a ComputeEngine protocol, any PyArrow logic outside the module boundary bypasses the protocol and prevents clean substitution.
Audit
Grepped pyiceberg/ (excluding io/pyarrow.py and tests) for runtime import pyarrow statements (both top-level and inline). Excluded TYPE_CHECKING-only imports since those have no runtime dependency.
| Location |
What it does |
Action |
table/upsert_util.py |
PyArrow table joins, group_by, compute, cast, take |
Absorb |
table/inspect.py |
Builds pa.schema + pa.Table.from_pylist for metadata inspection |
TBD |
transforms.py |
pyarrow_transform() dispatch on pa.Array/ChunkedArray |
TBD |
table/__init__.py |
Entry points accept pa.Table, delegate to io.pyarrow |
Leave |
table/deletion_vector.py |
Single pa.chunked_array() call |
Leave |
catalog/__init__.py |
Delegates to io.pyarrow for schema conversion |
Leave |
Plan
One PR per absorption. Each is a pure refactor: move code into io/pyarrow.py, have the caller import from pyiceberg.io.pyarrow instead of pyarrow directly. No behavior change, all existing tests pass unchanged.
Related
Summary
Before decomposing
pyiceberg/io/pyarrow.pyinto focused submodules (#3737, #3738), we should consolidate PyArrow-specific logic that currently lives outside the module. This ensures all PyArrow calls route through a single boundary, making the subsequent split clean and enabling future engine substitution.Motivation
Per discussion in #3737, @rambleraptor noted that the first useful step is ensuring no PyArrow logic occurs outside
pyarrow.py. Currently several modules importpyarrowdirectly and implement compute logic inline rather than delegating throughpyiceberg.io.pyarrow.When we later introduce a
ComputeEngineprotocol, any PyArrow logic outside the module boundary bypasses the protocol and prevents clean substitution.Audit
Grepped
pyiceberg/(excludingio/pyarrow.pyand tests) for runtimeimport pyarrowstatements (both top-level and inline). ExcludedTYPE_CHECKING-only imports since those have no runtime dependency.table/upsert_util.pytable/inspect.pytransforms.pypyarrow_transform()dispatch on pa.Array/ChunkedArraytable/__init__.pytable/deletion_vector.pycatalog/__init__.pyPlan
One PR per absorption. Each is a pure refactor: move code into
io/pyarrow.py, have the caller import frompyiceberg.io.pyarrowinstead ofpyarrowdirectly. No behavior change, all existing tests pass unchanged.table/upsert_util.pyPyArrow logictable/inspect.py(pending discussion)transforms.py(pending discussion)Related