Integration with PyTorch
PyTorch tensors can be built from DuckDB query results by exporting each result batch as Apache Arrow. This pattern is useful when training or serving models on Parquet data because DuckDB can filter and project the input before the batches are converted to tensors.
Installation
pip install -U duckdb pyarrow torchDuckDB to PyTorch
This example queries a Parquet-backed Arrow dataset, streams the result as RecordBatch objects, and converts each batch into feature and label tensors.
import pathlibimport tempfile
import duckdbimport numpy as npimport pyarrow as paimport pyarrow.dataset as dsimport pyarrow.parquet as pqimport torch
base_path = pathlib.Path(tempfile.mkdtemp())parquet_dir = base_path / "train"
table = pa.table( { "feature_0": [0.1, 0.3, 0.5, 0.7], "feature_1": [1.0, 0.0, 1.0, 0.0], "label": [0, 1, 1, 0], })pq.write_to_dataset(table, str(parquet_dir))
con = duckdb.connect()train_dataset = ds.dataset(str(parquet_dir))
reader = con.execute(""" SELECT feature_0, feature_1, label FROM train_dataset WHERE label = 1""").to_arrow_reader(batch_size=2)
for batch in reader: features = torch.tensor( np.column_stack( [ batch.column(0).to_numpy(), batch.column(1).to_numpy(), ] ), dtype=torch.float32, ) labels = torch.tensor(batch.column(2).to_numpy(), dtype=torch.int64)
print(features.shape, labels)torch.Size([2, 2]) tensor([1, 1])DuckDB pushes the WHERE label = 1 filter and the selected columns into the dataset scan, so only the requested rows and columns are converted to tensors. The to_arrow_reader method streams the result in batches instead of materializing the full result in memory at once.
To learn more about querying Arrow objects directly, see the “SQL on Apache Arrow” guide and the “Export to Apache Arrow” guide.