Install & Compatibility
Where this runs
tested against v2.0.10 · pip install
no network on importno background threads
Install × environment matrix
Each cell = how many times install + import succeeded across repeated harness runs. Partial = flaky.
glibc = Debian/Ubuntu slim · musl = Alpine Linux
muslpy 3.10–3.95 runs
build_error
glibcpy 3.10–3.95 runs
installs and imports cleanly · install 13.4s · import 0.646s · 354MB
358MB installed
● package 358MB
Code
Verified usage
Verified import paths — ran on the pinned version, not inferred.
create_engine
✓ from sqlalchemy import create_engine
DatabricksDialect (implicit)
✓ engine = create_engine('databricks://...')
The dialect is automatically registered with SQLAlchemy upon installation; no direct class import is typically needed. The 'databricks://' scheme is used in the connection URL.
TIMESTAMP
✓ from databricks.sqlalchemy import TIMESTAMP
For timezone-aware timestamps, as SQLAlchemy's default DateTime() is mapped to Databricks' timezone-agnostic TIMESTAMP_NTZ().
TIMESTAMP_NTZ
✓ from databricks.sqlalchemy import TIMESTAMP_NTZ
Explicitly import if you need the timezone-agnostic timestamp type.
TINYINT
✓ from databricks.sqlalchemy import TINYINT
Databricks-specific integer type.
This quickstart demonstrates how to establish a connection to Databricks using the `databricks-sqlalchemy` dialect and `create_engine`. It utilizes environment variables for secure authentication and connection details, then executes a simple SELECT query. It also includes an example of a parameterized query, highlighting the named paramstyle required by the dialect.
import os
from sqlalchemy import create_engine, text
# Ensure these environment variables are set:
# DATABRICKS_SERVER_HOSTNAME: Hostname of your Databricks workspace or SQL warehouse
# DATABRICKS_HTTP_PATH: HTTP path of your SQL warehouse or cluster
# DATABRICKS_TOKEN: Your Databricks Personal Access Token
# DATABRICKS_CATALOG (optional): Initial Unity Catalog catalog
# DATABRICKS_SCHEMA (optional): Initial Unity Catalog schema
host = os.environ.get('DATABRICKS_SERVER_HOSTNAME', 'your_databricks_host.cloud.databricks.com')
http_path = os.environ.get('DATABRICKS_HTTP_PATH', '/sql/1.0/endpoints/your_http_path')
access_token = os.environ.get('DATABRICKS_TOKEN', 'dapi...')
catalog = os.environ.get('DATABRICKS_CATALOG', 'default')
schema = os.environ.get('DATABRICKS_SCHEMA', 'default')
try:
connection_string = (
f"databricks://token:{access_token}@{host}?"
f"http_path={http_path}&catalog={catalog}&schema={schema}"
)
engine = create_engine(connection_string)
with engine.connect() as connection:
# Example: Execute a simple query
result = connection.execute(text("SELECT 1 AS one, 'hello' AS two"))
for row in result:
print(f"Row: {row.one}, {row.two}")
# Example with parameterized query (requires DBR 14.2+)
# Note: Databricks dialect uses 'named' paramstyle ':param'
param_value = "world"
result = connection.execute(text("SELECT :val AS greeting"), {'val': param_value})
for row in result:
print(f"Greeting: {row.greeting}")
print("Successfully connected and executed queries.")
except Exception as e:
print(f"Error connecting to Databricks: {e}")
print("Please ensure environment variables are correctly set and Databricks resources are accessible.")
Debug
Known issues
breakingThe `databricks-sqlalchemy` library has distinct major versions for SQLAlchemy 1.x and SQLAlchemy 2.x. Version `1.x` of `databricks-sqlalchemy` supports SQLAlchemy 1.x, while `2.x` supports SQLAlchemy 2.x. Upgrading `databricks-sqlalchemy` to `2.x` requires a corresponding upgrade of `SQLAlchemy` to `2.x`, and vice-versa, to ensure compatibility.fixEnsure `databricks-sqlalchemy` and `SQLAlchemy` major versions are compatible. For SQLAlchemy 2.x, use `pip install databricks-sqlalchemy`. For SQLAlchemy 1.x, use `pip install databricks-sqlalchemy~=1.0`.
affects: All versions (specifically, the distinction between 1.x and 2.x branches)
gotchaThere are two similarly named, but distinct, Python packages: `databricks-sqlalchemy` (the official dialect, this package) and `sqlalchemy-databricks` (an older, separate community-driven package). Always ensure you install and use `databricks-sqlalchemy` for official and up-to-date support.fixVerify your `pip install` command is `pip install databricks-sqlalchemy` and check your `requirements.txt` to avoid `sqlalchemy-databricks`.
affects: All versions
gotchaThe SQLAlchemy 2.0 dialect for Databricks *always* uses native parameterized queries with the 'named' paramstyle (`:param`) and requires Databricks Runtime 14.2 or above. Using a different paramstyle or older DBR versions will result in query failures.fixEnsure your Databricks Runtime is 14.2 or newer, and use the `:param_name` syntax for parameters in your SQL queries when executing via SQLAlchemy.
affects: 2.0.0 and later
gotchaSQLAlchemy's `DateTime()` type is mapped to Databricks' timezone-agnostic `TIMESTAMP_NTZ()`. If you declare `DateTime(timezone=True)`, the `timezone` argument will be ignored by the Databricks dialect. For timezone-aware timestamps, you must explicitly import and use `databricks.sqlalchemy.TIMESTAMP()`.fixUse `databricks.sqlalchemy.TIMESTAMP()` for timezone-aware columns. For timezone-agnostic columns, `sqlalchemy.DateTime()` will correctly map to `TIMESTAMP_NTZ()`.
affects: All versions
gotchaOlder versions of `databricks-sqlalchemy` (e.g., 1.0.5) had issues with unquoted identifiers in Unity Catalog, particularly for catalog or schema names containing hyphens. This could lead to `[INVALID_IDENTIFIER]` errors during metadata operations.fixUpgrade to `databricks-sqlalchemy` version 2.x to resolve identifier quoting issues in Unity Catalog. Always quote identifiers in your raw SQL if they contain special characters.
affects: <2.0.0
gotchaThe `autoincrement` argument on SQLAlchemy type declarations (e.g., `Integer(autoincrement=True)`) is currently ignored by the Databricks dialect. To create an auto-incrementing field, you must explicitly use `sqlalchemy.Identity()` and combine it with the `databricks.sqlalchemy.BigInteger()` type, as only BIGINT fields support auto-increment in Databricks Runtime.fixFor auto-incrementing fields, define them as `Column(BigInteger, Identity())`.
affects: All versions
breakingInstallation in minimal environments (e.g., Alpine Linux) may fail for packages with C extensions if essential build tools are missing. Many Python packages, including some dependencies of `databricks-sqlalchemy`, require a C compiler and Python development headers to be built from source during installation.fixEnsure your environment includes necessary build tools. For Alpine Linux, install `build-base` and `python3-dev` using `apk add build-base python3-dev`. For Debian/Ubuntu, use `apt-get install build-essential python3-dev`.
affects: All versions
Errors
Common errors & fixes
sqlalchemy.exc.NoSuchModuleError: Can't load plugin: sqlalchemy.dialects:databricks
This error occurs when the SQLAlchemy library cannot find or load the Databricks dialect because the `databricks-sqlalchemy` package is not correctly installed or registered with SQLAlchemy.
fixEnsure `databricks-sqlalchemy` is installed in your environment: `pip install databricks-sqlalchemy`. Also, verify that SQLAlchemy itself is installed and up-to-date.
ModuleNotFoundError: No module named 'databricks'
This typically means that the `databricks-sqlalchemy` package or its core dependency, `databricks-sql-connector`, has not been installed in the Python environment.
fixInstall the library using pip: `pip install databricks-sqlalchemy`. If `databricks-sql-connector` is missing, install it explicitly: `pip install databricks-sql-connector`.
AttributeError: module 'sqlalchemy.types' has no attribute 'Uuid'
This error usually indicates a version incompatibility between SQLAlchemy and a dependent library (like `langchain`) or between `databricks-sqlalchemy` and the installed SQLAlchemy version, where the 'Uuid' type is expected but not found in the current SQLAlchemy version.
fixUpgrade SQLAlchemy to version 2.0 or higher: `pip install --upgrade SQLAlchemy` or specifically to a compatible version that includes the `Uuid` type if you are using an older version of SQLAlchemy or related libraries that require it.
sqlalchemy.exc.OperationalError: (databricks.sql.exc.RequestError) Error during request to server: BAD_REQUEST: Parameterized query has too many parameters: 2000 parameters were given but the limit is 256.
This specific `OperationalError` occurs when attempting to insert too many rows or columns at once using methods like `pandas.to_sql` with `method='multi'`, exceeding the maximum parameter limit enforced by the Databricks SQL endpoint.
fixWhen using `pandas.to_sql`, avoid setting `method='multi'` or break down the DataFrame into smaller chunks before inserting to stay within the parameter limit. Alternatively, try inserting row-by-row, though this might be slower.
Upgrade
Version history
2.0.10latest on PyPI · released Jul 2, 2026
Audit
Dependencies
sqlalchemyrequiredCore ORM and SQL toolkit that this library extends.
databricks-sql-connectorrequiredThe underlying Python driver used by the dialect; SQLAlchemy features were moved from this connector into databricks-sqlalchemy in v4.0.0+ of the connector. It is crucial for functionality.
pyarrowoptionalOptional dependency for enhanced performance with large data volumes, enabling features like CloudFetch in the underlying databricks-sql-connector.