Registry / gcp / google-cloud-bigquery-storage

google-cloud-bigquery-storage

JSON →
library2.41.0pypypi✓ verified 24d ago

The Google Cloud BigQuery Storage API client library enables high-throughput data transfer from BigQuery tables. It leverages a binary serialization format (like Apache Arrow or Protobuf) for efficient data transfer and is ideal for analytical workloads requiring large-scale data extraction. The library is currently at version 2.36.2 and is released as part of the `google-cloud-python` monorepo, following a frequent release cadence.

pip install google-cloud-bigquery-storage
INSTALL
IMPORT
SIG · GOOGLE-CLOUD-BIGQU
G
google-cloud-bigquery-storage
gcppythonv2.41.0
Install
10.8s avg
Import
2041ms
Disk
69MB
Pass rate
10/ 10
Env Coverage10 / 10
glibc
3.93.13
musl
3.93.13
Install & Compatibility
Where this runs
tested against v2.41.0 · pip install
no network on importno background threads
Install × environment matrix
Each cell = how many times install + import succeeded across repeated harness runs. Partial = flaky.
glibc = Debian/Ubuntu slim · musl = Alpine Linux
musl
py 3.103.910 runs
installs and imports cleanly · install 0.0s · import 2.298s · 71.1MB
glibc
py 3.103.910 runs
installs and imports cleanly · install 10.8s · import 1.784s · 69MB
69MB installed
● package 69MB
Code
Verified usage

Verified import paths — ran on the pinned version, not inferred.

BigQueryReadClient
from google.cloud.bigquery_storage import BigQueryReadClient
from google.cloud.bigquery_storage_v1 import BigQueryReadClient
The canonical import path changed in version 2.0.0. The `_v1` suffix is no longer recommended.
BigQueryWriteClient
from google.cloud.bigquery_storage import BigQueryWriteClient
from google.cloud.bigquery_storage_v1 import BigQueryWriteClient
The canonical import path changed in version 2.0.0. The `_v1` suffix is no longer recommended.
types
from google.cloud.bigquery_storage import types
from google.cloud.bigquery_storage_v1 import enums
Enum types are now accessed via the `types` module (e.g., `types.DataFormat.ARROW`). Direct import from `enums` or via the client (`BigQueryReadClient.enums`) is deprecated since v2.0.0.

This quickstart demonstrates how to use `BigQueryReadClient` to read data from a public BigQuery table using the Storage API. It configures a read session, reads data in Apache Arrow format, and attempts to convert it to a Pandas DataFrame. Remember to set the `GOOGLE_CLOUD_PROJECT` environment variable and ensure the BigQuery Storage API is enabled for your project.

import os from google.cloud.bigquery_storage import BigQueryReadClient, types # Your Google Cloud project ID. If not set, it will default to the project # defined in your environment (e.g., GOOGLE_CLOUD_PROJECT, gcloud config). project_id = os.environ.get('GOOGLE_CLOUD_PROJECT', '') if not project_id: raise ValueError("GOOGLE_CLOUD_PROJECT environment variable must be set") # Public BigQuery dataset and table to read from parent = f"projects/{project_id}" dataset_id = "google_trends" table_id = "international_top_rising_terms" table = f"projects/bigquery-public-data/datasets/{dataset_id}/tables/{table_id}" def read_bigquery_table_storage(project_id, table): """Reads data from a BigQuery table using the BigQuery Storage Read API.""" client = BigQueryReadClient() # Specify the table and desired data format (Arrow recommended for performance) read_options = types.ReadSession.TableReadOptions(selected_fields=["country_name", "region_name"]) requested_session = types.ReadSession( table=table, data_format=types.DataFormat.ARROW, # Or types.DataFormat.AVRO read_options=read_options, ) # Create a read session read_session = client.create_read_session( parent=parent, read_session=requested_session, max_stream_count=1 # Adjust for parallelism if needed ) print(f"Read session created: {read_session.name}") # Read from the first stream (assuming max_stream_count=1) stream_name = read_session.streams[0].name reader = client.read_rows(stream_name) # Convert to Pandas DataFrame (requires pandas and pyarrow installed) import pandas as pd try: dataframe = reader.to_dataframe() print("Successfully read data into Pandas DataFrame.") print(dataframe.head()) return dataframe except ImportError as e: print(f"Could not convert to DataFrame: {e}. Try iterating rows manually.") print("Reading rows directly:") for row_message in reader.rows(session=read_session): # row_message will be a dict if fastavro is installed and format is AVRO, # or a protobuf message if format is ARROW (requires manual parsing for dicts) print(row_message) break # Print first row and exit if __name__ == "__main__": # Ensure you have authenticated to GCP (e.g., `gcloud auth application-default login`) # and enabled the BigQuery Storage API for your project. # Replace 'your-gcp-project-id' with your actual project ID or set GOOGLE_CLOUD_PROJECT env var. # Example uses a public dataset, so you mainly need read access to your billing project. df = read_bigquery_table_storage(project_id, table)
Debug
Known issues
breakingMajor breaking changes occurred in version 2.0.0. The primary import path for clients changed from `google.cloud.bigquery_storage_v1` to `google.cloud.bigquery_storage`. Enum types moved from direct import or client access (e.g., `BigQueryReadClient.enums`) to the `types` module (e.g., `types.DataFormat.ARROW`). Existing code using the `_v1` suffix or direct enum access will fail.
fix
Update import statements to `from google.cloud.bigquery_storage import ...` and access enums via `from google.cloud.bigquery_storage import types`.
affects: >=2.0.0
deprecatedThe `client_config` and `channel` parameters for client constructors have been removed.
fix
Remove these parameters from client instantiation. Customize retry and timeout settings directly when invoking methods (e.g., `client.create_read_session(..., timeout=60)`).
affects: >=2.0.0
gotchaThe BigQuery Storage API is optimized for high-throughput, large-scale data transfer, not for small, interactive queries or single-row lookups. Using it for small datasets may introduce unnecessary overhead compared to the standard BigQuery client library.
fix
Use the `google-cloud-bigquery` client library for standard SQL queries, small data extractions, or metadata operations. Reserve `google-cloud-bigquery-storage` for high-volume data ingestion or extraction workflows.
affects: All
gotchaData read from the Storage API is returned in a binary format (Protobuf or Apache Arrow). To easily work with this data in Python, you typically need to convert it. This often requires additional dependencies like `pandas` and `pyarrow` (for `to_dataframe()`) or `fastavro` (for `rows()` to get dicts from AVRO).
fix
Install `google-cloud-bigquery-storage` with appropriate extras (e.g., `pip install google-cloud-bigquery-storage[fastavro,pandas,pyarrow]`) and use methods like `reader.to_dataframe()` or `reader.rows()`.
affects: All
gotchaWhile `BigQueryReadClient.create_read_session` allows specifying `max_stream_count` for parallelism, achieving true concurrent data processing in Python often requires using the `multiprocessing` module rather than simple threading due to Python's Global Interpreter Lock (GIL).
fix
For optimal parallel performance when reading multiple streams, consider using Python's `multiprocessing` module or an asynchronous framework if your I/O operations are truly non-blocking across network requests.
affects: All
gotchaThe `google-cloud-bigquery` client library, which often complements `google-cloud-bigquery-storage`, has ended support for Python 3.7 and 3.8. Although `google-cloud-bigquery-storage` officially supports Python >=3.7, it is highly recommended to upgrade to Python 3.9+ to maintain compatibility with the broader Google Cloud client ecosystem and ensure ongoing support.
fix
Upgrade your Python environment to 3.9 or higher.
affects: All (especially for Python 3.7, 3.8 users)
gotchaMost Google Cloud client libraries, including `google-cloud-bigquery-storage`, require proper authentication. In many development environments, this is achieved by setting the `GOOGLE_CLOUD_PROJECT` environment variable to specify the project ID, or by providing explicit credentials (e.g., via `GOOGLE_APPLICATION_CREDENTIALS`). Failure to do so will result in authentication errors or errors like 'GOOGLE_CLOUD_PROJECT environment variable must be set'.
fix
Ensure the `GOOGLE_CLOUD_PROJECT` environment variable is set to your Google Cloud project ID (e.g., `export GOOGLE_CLOUD_PROJECT='your-project-id'`) before running your application. Alternatively, explicitly pass project credentials to the client constructor, or ensure `GOOGLE_APPLICATION_CREDENTIALS` points to a valid service account key file.
affects: All
gotchaGoogle Cloud client libraries often require a project ID for operations. If not explicitly provided during client instantiation or via default credentials (e.g., service account JSON), the library typically looks for the `GOOGLE_CLOUD_PROJECT` environment variable. Failure to set this will result in a `ValueError`.
fix
Ensure the `GOOGLE_CLOUD_PROJECT` environment variable is set to your Google Cloud project ID, or explicitly pass the `project` argument to the client constructor (e.g., `BigQueryReadClient(project='your-project-id')`).
affects: All
Errors
Common errors & fixes
ModuleNotFoundError: No module named 'google.cloud.bigquery_storage'
The `google-cloud-bigquery-storage` library is not installed in the Python environment, or the Python interpreter cannot find it in its path.
fix
Install the library using pip: `pip install google-cloud-bigquery-storage`
ImportError: cannot import name 'bigquery_storage_v1beta1' from 'google.cloud'
This error typically occurs when trying to import an older, deprecated version of the BigQuery Storage API client (`v1beta1`) after upgrading the `google-cloud-bigquery-storage` library to a newer version (2.x or later), which primarily uses the `v1` or top-level `bigquery_storage` namespace.
fix
Update your import statements to use the current API version or the top-level namespace: `from google.cloud.bigquery_storage import BigQueryReadClient, types` or `from google.cloud import bigquery_storage`
ValueError: The pyarrow library is not installed, please install pyarrow to use the to_arrow() function.
The `google-cloud-bigquery-storage` library relies on `pyarrow` for efficient data serialization when working with Apache Arrow format or converting results to Pandas DataFrames, but `pyarrow` is not installed as a dependency.
fix
Install the `pyarrow` library: `pip install pyarrow` or install with the BigQuery client extra: `pip install google-cloud-bigquery[bqstorage,pandas]`
AttributeError: 'NoneType' object has no attribute '_parse_avro_schema'
This error can occur when using `reader.to_dataframe()` on a `ReadRowsIterable` object that returns no rows, or when there's an issue parsing the Avro schema, sometimes related to specific data types or filters on the table.
fix
Ensure that the query or read session is expected to return data. If no data is expected, handle the case where the `ReadRowsIterable` might be empty before attempting to convert to a DataFrame. This might involve checking for data presence or wrapping the conversion in a `try-except` block. Additionally, verify `fastavro` is installed as it's required for Avro serialization, `pip install fastavro`.
Upgrade
Version history
2.41.0latest on PyPI · released Aug 24, 2026
Audit
Dependencies
pyarrowoptionalRequired for `to_arrow()` method to read data in Apache Arrow format, which is highly efficient.
pandasoptionalRequired for `to_dataframe()` method to convert read data directly into a Pandas DataFrame. Requires `pyarrow` or `fastavro` as well.
fastavrooptionalRequired for `rows()` method to parse protobuf messages into Python dictionaries when not using Arrow format. Useful when `pandas` or `pyarrow` are not installed.
google-cloud-bigqueryoptionalOften used in conjunction to manage BigQuery tables, run standard SQL queries, or perform operations not related to high-throughput data transfer. Not a direct dependency for `bigquery-storage` itself.
Agent activity
27 hits · last 30 days
node
22
OpenAI (training)
1
Resources